Publications 2026

This list includes all the publications which were created in the context of the LLM4DH project in 2026.

Book chapters

Journal articles

  • Ulčar, M., Žagar, A., Armendariz, C.S., Repar, A., Pollak, S., Purver, M., and Robnik Šikonja, M. (2026). Mono- and cross-lingual evaluation of representation language models on less-resourced languages, Computer Speech & Language, 95, 101852. https://doi.org/10.1016/j.csl.2025.101852
  • Verdonik, D., and Vidinić, J. (2026). Dialoška dejanja v zasebni govorni interakciji / Dialogue acts in private spoken interaction. Slavistična revija, 74(1), 77–96. https://srl.si/ojs/srl/article/view/4295

Conference papers

  • Arčon, T. (2026). Lost in instructions? Developing a systematic approach to instruction tuning datasets for LLMs (Tjaša Arčon). Proceedings of the 17th Postgraduate International Conference on Advancements in Computer Science and Applications PICACSA, pp. 1-9.
  • Arhar Holdt, Š. & I. Kosem. (2026). Lifelong Development of Literacy Skills. Conference on Literacy (6. – 8. July 2026). https://www.literacyeurope.org/wp-content/uploads/2026/07/Abstract-Book-Zbirka-povzetkov.pdf
  • Blevins, T., Mayhew, S., Suppa, M., Gonen, H., Mirkin, S., Pais, V., Dobrovoljc, K., Giouli, V., Kevin, J., Jang, E., Kim, E., Seo, J., Gialis, X., & Pinter, Y. (2026). Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) (pp. 6609–6618). European Language Resources Association (ELRA). https://doi.org/10.63317/4qhcjikvgeda.
  • Caporusso, J., Hoogland, D., Koloski, B., Purver, M., Pollak, S., & Vintar, S. (2026). Exploring Social Bias in Slovenia: The EEC-SL Dataset. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) (pp. 4019–4030). European Language Resources Association (ELRA). https://doi.org/10.63317/2pdt2x4ci6e5.
  • Dobrovoljc, K., Verdonik, D., Čibej, J., Rupnik, P., & Ljubešić, N. (2026). ROG: A Multi-Layer Manually Annotated Corpus of Spoken Slovenian. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) (pp. 5701–5710). European Language Resources Association (ELRA). https://doi.org/10.63317/44np3dvumisj.
  • Klemen, M., Arčon, T., Terčon, L., Robnik-Šikonja, M., & Dobrovoljc, K. (2026). Towards corpus-grounded agentic LLMs for multilingual grammatical analysis. In E. Hinrichs, J. Nivre, P. Osenova, J. Pustejovsky, & C. Zinn (Eds.), Proceedings of the Workshop on Structured Linguistic Data and Evaluation (SLiDE) (pp. 136–147).
  • Knez, T. and Žitnik, S. (2026). Improving Slovene Language Models for Lexicographic Question Answering through Continued Pretraining and Instruction Fine-Tuning. Proceedings of the Workshop on Structured Linguistic Data and Evaluation, pp. 114–123. https://www.slide-workshop.org/book.pdf#page=128
  • Martinc, M., & Vreš, D. (2026). A Large-Scale Instruction-Tuning Dataset and Models for Slovenian Vision-Language Tasks. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) (pp. 9420–9433). European Language Resources Association (ELRA). https://doi.org/10.63317/2e3jf6e7tcoh.
  • Pannitto, L., Dobrovoljc, K., & Guillaume, B. (2026). Survey of Tools for Manual Linguistic Annotation: Supporting Diversity through Interactive Exploration. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) (pp. 11562–11573). European Language Resources Association (ELRA). https://doi.org/10.63317/4u4be6bbtj8e.
  • Robnik-Šikonja, M. (2026). Text readability improvement with large language models. Conference on Literacy (6. – 8. July 2026).
  • Štebljaj, Ž. (2026). Generating Grammatical Errors Explanations for Slovene. Proceedings of the 17th Postgraduate International Conference on Advancements in Computer Science and Applications PICACSA, pp. 193-200.
  • Vreš, D. (2026). Robust LLM for a Less-Resourced Language. Proceedings of the 17th Postgraduate International Conference on Advancements in Computer Science and Applications PICACSA, pp. 209-216.
  • Zgank, A., Donaj, G., Kolaric, U., Sereinig, U., Koren-Zwitter, T., Boto, S., Zwitter-Grilc, S., Vidinic, J., & Verdonik, D. (2026). Developing Zila: A Spoken Language Resource for the Endangered Slovenian Gail Valley Dialect. In Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) (pp. 3325–3332). European Language Resources Association (ELRA). https://doi.org/10.63317/3fhhk948chhm.

Datasets

  • Arčon, T.; Klemen, M.; Robnik-Šikonja, M.; Dobrovoljc, K. and Terčon, L. (2026). A multilingual benchmark for evaluating metalinguistic knowledge WALS-Bench 1.0, Slovenian language resource repository CLARIN.SI.
  • Kuzman Pungeršek, T., Rupnik, P. and Ljubešić, N. (2026). South Slavic web corpus collection CLASSLA-web 2.0, Slovenian language resource repository CLARIN.SI. http://hdl.handle.net/11356/2079.
  • Terčon, L.; Dobrovoljc, K.; Klemen, M., Arčon, T. and Robnik-Šikonja, M. (2026). Corpus-grounded evaluation dataset for grammatical question answering GramQA 1.0, Slovenian language resource repository CLARIN.SI. http://hdl.handle.net/11356/2086.

Other

Large Language Models for Digital Humanities (2026). ContRAG: Contradiction-Aware Retrieval for Legal Texts [Large Language Model]. LLM4DH. https://github.com/clarinsi/LegalContradictionRAG