{"id":1699,"date":"2026-08-03T10:07:03","date_gmt":"2026-08-03T10:07:03","guid":{"rendered":"https:\/\/intellexus.net\/?p=1699"},"modified":"2026-08-03T10:07:03","modified_gmt":"2026-08-03T10:07:03","slug":"id10m-jam-stress-testing-idiom-identification-under-challenging-context","status":"publish","type":"post","link":"https:\/\/intellexus.net\/index.php\/2026\/08\/03\/id10m-jam-stress-testing-idiom-identification-under-challenging-context\/","title":{"rendered":"\u201cID10M-JAM: Stress-Testing Idiom Identification Under Challenging Context\u201d"},"content":{"rendered":"\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\"><a href=\"https:\/\/aclanthology.org\/people\/kai-golan-hashiloni\/\">Kai Golan Hashiloni<\/a>,\u00a0<a href=\"https:\/\/aclanthology.org\/people\/lior-livyatan\/unverified\/\">Lior Livyatan<\/a>,\u00a0<a href=\"https:\/\/aclanthology.org\/people\/ofri-hefetz\/unverified\/\">Ofri Hefetz<\/a>,\u00a0<a href=\"https:\/\/aclanthology.org\/people\/alon-mannor\/unverified\/\">Alon Mannor<\/a>,\u00a0<a href=\"https:\/\/aclanthology.org\/people\/bar-cohen\/unverified\/\">Bar Cohen<\/a>,\u00a0<a href=\"https:\/\/aclanthology.org\/people\/kfir-bar\/\">Kfir Bar<\/a>, 2026.<\/p>\n\n\n\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\">Large language models (LLMs) achieve strong performance on idiom identification benchmarks, yet their robustness to misleading contextual signals remains largely untested. We introduce ID10M-JAM, an adversarial extension of the ID10M dataset designed to jam model understanding by injecting coherent but conflicting context before each target sentence. For every sentence containing a potential idiomatic expression (PIE), we construct variants that deliberately invert contextual expectations: placing literal cues before idiomatic uses and idiomatic cues before literal ones. All variants are validated by human annotators to ensure naturalness and unambiguous interpretation for human readers. ID10M-JAM exposes systematic vulnerabilities in LLMs\u2019 contextual reasoning, pushing idiom identification to its breaking point.<\/p>\n\n\n\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\"><br><a href=\"https:\/\/aclanthology.org\/2026.findings-acl.1045\/\">https:\/\/aclanthology.org\/2026.findings-acl.1045\/<\/a><\/p>\n\n\n\n<p style=\"font-size:clamp(0.929rem, 0.929rem + ((1vw - 0.2rem) * 0.785), 1.4rem);\"><\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/intellexus.net\/wp-content\/uploads\/2026\/03\/DharmaBench_2025.ijcnlp-long.114.pdf\">Download PDF<\/a><\/div>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Kai Golan Hashiloni,\u00a0Lior Livyatan,\u00a0Ofri Hefetz,\u00a0Alon Mannor,\u00a0Bar Cohen,\u00a0Kfir Bar, 2026. Large language models (LLMs) achieve strong performance on idiom identification benchmarks, yet their robustness to misleading contextual signals remains largely untested. We introduce ID10M-JAM, an adversarial extension of the ID10M dataset designed to jam model understanding by injecting coherent but conflicting context before each target sentence. [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[21],"tags":[],"class_list":["post-1699","post","type-post","status-publish","format-standard","hentry","category-publications"],"_links":{"self":[{"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/posts\/1699","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/comments?post=1699"}],"version-history":[{"count":2,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/posts\/1699\/revisions"}],"predecessor-version":[{"id":1701,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/posts\/1699\/revisions\/1701"}],"wp:attachment":[{"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/media?parent=1699"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/categories?post=1699"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/intellexus.net\/index.php\/wp-json\/wp\/v2\/tags?post=1699"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}