Geology of Texts, Genealogy of Concepts, Intellectual Ecosystems: Mapping the Indic and Tibetic Buddhist Text Corpora

ERC Synergy Grant Project no. 101118558 (April 2024 – March 2030)

India and Tibet have both been home to great Buddhist cultures, producing a vast quantity of learned literature in respectively Sanskrit and Tibetan. In recent years, a flood of such texts has become available, also in digital form.

Specialists have long studied Sanskrit and Tibetan Buddhist texts with painstaking philological-historical methods. Now, inspired by the profusion of materials that has become available, and recent advances in computer sciences (Artificial Intelligence and Natural Language Processing in particular), a team of experienced textual scholars from the University of Hamburg has joined forces with specialists in computer sciences from Reichman University, to carry out an innovative project which aims to develop cutting-edge computational tools to revolutionize the study of this material. Chief among these are the automatic identification and matching of concepts, both mono- and cross-lingually, tasks that go far beyond the matching of texts.

In addition to developing computational models and tools, which have the potential to bring about a transformation also in the study of similar materials from other cultures, the project will map three large Buddhist text corpora in Sanskrit and Tibetan, shedding light on their interdependence, their interactions also with other, non-Buddhist, texts, and the processes of their evolution.

  • “MoraLink: Bridging Morals and Narrative Fables for Retrieval”

    Lior Livyatan, Kai Golan Hashiloni, Asif Amar, Roei Rahamim, Kfir Bar, 2026. A fable is a short narrative whose purpose is to convey a moral lesson. While semantically related, morals and fables differ fundamentally in form: morals are abstract and propositional, while fables express the same idea through characters and plot. In this work, we…

  • “IdioSteer: Controlling Figurative-vs-Literal Interpretation via Residual-Stream Activation Steering”

    Bar Cohen, Kai Golan Hashiloni, Kfir Bar, 2026. Large language models (LLMs) default to figurative interpretations of idiomatic expressions when context supports them. We investigate whether this bias can be directly controlled via activation steering: asking whether idiomaticity is encoded as a manipulable direction in residual space. We introduce IdioSteer, a small controlled benchmark of…

  • “IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions”

    Kai Golan Hashiloni, Daniel Fadlon, Lior Livyatan, Ofri Hefetz, Jiahuan Pei, Kfir Bar, 2026. Idioms pose a fundamental challenge for language models, as their meaning cannot be inferred from surface form alone. Understanding such expressions, therefore, requires semantic abstraction beyond lexical overlap. We introduce IdioLink, a retrieval benchmark designed to test whether models can link…