Publications

“Scaling Sentence Similarity for Classical Tibetan with Automatic Annotations”


Shay Cohen, Jingyi Yang, Gal Rabinovitz, Sonam Choden, Ofir Shtrosberg, Nicola Bajetta, Goody Ben Horin, Rebecca Sundén, Omri Drori, Sonam Jamtsho, Dorji Wangchuk, Kfir Bar, Orna Almogi, Shai Fine, 2026.

Identifying intertextual parallels is a central task in textual philology, traditionally requiring labor-intensive manual analysis. While digitized historical corpora enable automated approaches using semantic sentence embeddings, training such models requires large annotated datasets, which are scarce for low-resource historical languages.

We address this challenge by introducing a scalable, automated annotation pipeline to train semantic embedding models for Classical Tibetan. Our method combines unsupervised contrastive bootstrapping with iterative pair mining, generating silver-standard similarity labels through two complementary annotation strategies: (1) a representation ensemble of embedding models and rerankers, and (2) an LLM-as-a-judge committee using best–worst scaling. When combined with a domain-specific gold dataset for sequential fine-tuning, the resulting model achieves a state-of-the-art Spearman correlation of 0.864, enabling effective semantic search in Classical Tibetan and offering a general framework for synthetic supervision in low-resource digital humanities settings.


https://aclanthology.org/2026.nlp4dh-1.15/