{"id":1713,"date":"2026-08-10T09:54:34","date_gmt":"2026-08-10T09:54:34","guid":{"rendered":"https:\/\/intellexus.net\/?p=1713"},"modified":"2026-08-10T09:54:34","modified_gmt":"2026-08-10T09:54:34","slug":"scaling-sentence-similarity-for-classical-tibetan-with-automatic-annotations","status":"publish","type":"post","link":"https:\/\/intellexus.net\/index.php\/2026\/08\/10\/scaling-sentence-similarity-for-classical-tibetan-with-automatic-annotations\/","title":{"rendered":"\u201cScaling Sentence Similarity for Classical Tibetan with Automatic Annotations\u201d"},"content":{"rendered":"\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\">Shay Cohen,\u00a0Jingyi Yang,\u00a0Gal Rabinovitz,\u00a0Sonam Choden,\u00a0Ofir Shtrosberg,\u00a0Nicola Bajetta,\u00a0Goody Ben Horin,\u00a0Rebecca Sund\u00e9n,\u00a0Omri Drori,\u00a0Sonam Jamtsho,\u00a0Dorji Wangchuk,\u00a0Kfir Bar,\u00a0Orna Almogi,\u00a0Shai Fine, 2026.<\/p>\n\n\n\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\">Identifying intertextual parallels is a central task in textual philology, traditionally requiring labor-intensive manual analysis. While digitized historical corpora enable automated approaches using semantic sentence embeddings, training such models requires large annotated datasets, which are scarce for low-resource historical languages.<\/p>\n\n\n\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\">We address this challenge by introducing a scalable, automated annotation pipeline to train semantic embedding models for Classical Tibetan. Our method combines unsupervised contrastive bootstrapping with iterative pair mining, generating silver-standard similarity labels through two complementary annotation strategies: (1) a representation ensemble of embedding models and rerankers, and (2) an LLM-as-a-judge committee using best\u2013worst scaling. When combined with a domain-specific gold dataset for sequential fine-tuning, the resulting model achieves a state-of-the-art Spearman correlation of 0.864, enabling effective semantic search in Classical Tibetan and offering a general framework for synthetic supervision in low-resource digital humanities settings.<\/p>\n\n\n\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\"><br><a href=\"https:\/\/aclanthology.org\/2026.nlp4dh-1.15\/\">https:\/\/aclanthology.org\/2026.nlp4dh-1.15\/<\/a><\/p>\n\n\n\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\"><\/p>\n\n\n\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\"><\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/intellexus.net\/wp-content\/uploads\/2026\/08\/2026.nlp4dh-1.15-Shay-Cohen.pdf\">Download PDF<\/a><\/div>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Shay Cohen,\u00a0Jingyi Yang,\u00a0Gal Rabinovitz,\u00a0Sonam Choden,\u00a0Ofir Shtrosberg,\u00a0Nicola Bajetta,\u00a0Goody Ben Horin,\u00a0Rebecca Sund\u00e9n,\u00a0Omri Drori,\u00a0Sonam Jamtsho,\u00a0Dorji Wangchuk,\u00a0Kfir Bar,\u00a0Orna Almogi,\u00a0Shai Fine, 2026. Identifying intertextual parallels is a central task in textual philology, traditionally requiring labor-intensive manual analysis. While digitized historical corpora enable automated approaches using semantic sentence embeddings, training such models requires large annotated datasets, which are scarce for low-resource [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[21],"tags":[],"class_list":["post-1713","post","type-post","status-publish","format-standard","hentry","category-publications"],"_links":{"self":[{"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/posts\/1713","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/comments?post=1713"}],"version-history":[{"count":2,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/posts\/1713\/revisions"}],"predecessor-version":[{"id":1715,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/posts\/1713\/revisions\/1715"}],"wp:attachment":[{"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/media?parent=1713"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/categories?post=1713"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/tags?post=1713"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}