Data and Benchmarking
Research-ready datasets and benchmarks for interconnected living texts.
Learn moreNews
Four papers at EMNLP-2026. We are excited to share four papers from our colleagues at UKP Lab that have been accepted to EMNLP-2026! Zyska et al. present Exposía, a dataset connecting academic writing with peer and instructor feedback for teaching and assessing scientific writing skills. Purkayastha et al. propose an LLM-guided framework for reviewer feedback that identifies guideline violations and provides targeted, actionable suggestions to improve peer reviews. Basch et al. introduce ReGround, a large-scale benchmark for grounding reviewer comments in multimodal evidence from scientific papers. Kuznetsov et al. propose an evidence-based framework for measuring peer review quality across venues and time. Check out the papers to learn more!
Three papers at ACL-2026. We are excited to share three papers from our colleagues at UKP Lab that have been accepted to ACL-2026! Ruan et al. introduce Re³Align, REspGen, and REspEval for author-in-the-loop response generation and evaluation in peer review. Şahinuç et al. propose reward models for scientific writing evaluation that generalize across diverse tasks and evaluation criteria. Baumgärtner et al. present SciCoQA, a benchmark for detecting discrepancies between scientific papers and their code. Check out the papers to learn more!
Three papers at EACL-2026. We are excited to share three papers from our colleagues at UKP Lab that have been accepted to EACL-2026! Sukannya et al. model meta-reviewing as a document-grounded dialogue for deliberative decision-making. Serwar et al. introduce ABCD-LINK, a framework for bootstrapping fine-grained cross-document links. Osama et al. present an LLM-assisted approach for more rigorous and evidence-based novelty assessment in peer review. Check out the papers to learn more!
🚀 NLPeer v.2 expanded. One of the largest openly available peer review datasets now gets even larger: we are releasing an extension of the NLPeer v.2 dataset, now including ACL 2025 data. This update adds 2k papers, 2k reviews, 1.5k rebuttals, and 849 meta-reviews. Learn more about the project here, or go straight to the dataset!
Two papers at ACL. We are exited to share two papers from our colleagues at UKP Lab that have been accepted to ACL-2025! Nils et al. present a novel framework for text assessment, modeling it as a Natural Language Diagnostic Abductive Reasoning process. Sukannya et al. introduce a new dataset of peer-review sentences annotated with fine-grained categories of lazy thinking. Check out the papers to learn more!
STRICTA at ACL. Our paper on Natural Language Diagnostic Abductive Reasoning to appear in ACL-2025 Main! Have a look at the preprint and if you also go to Vienna, come to our presentation to meet the authors and talk about the work.
Three papers at NAACL. Intertextuality has many applications, and we are excited to share three InterText-related papers by our colleagues at UKP Lab, to appear at NAACL-2025! Jonathan Tonglet et al. address the problem of debunking misinformation in images. Max Glockner et al. investigate misrepresentation of scientific claims. And Tim Baumgärtner et al. explore question answering for academic peer reviews, winning an Outstanding Paper Award 🏆! Have a look at the preprints, and if you are at NAACL, visit their talks to learn more.
Keynote at SIGIR-25. We are happy to announce that Iryna Gurevych will give a Keynote on the use of AI for science and expert-AI collaboration at SIGIR-2025. If you are at the conference, come to our talk to learn more about what InterText has been up to in the past months, and about other related initiatives at UKP Lab.
🚀 NLPeer v.2 has arrived. After many months of hard work, we are happy to announce the release of the NLPeer v.2. corpus – a new iteration of the data collection initiative from ACL Rolling Review, ELIFE, PLOS and other venues. With over 1.8k papers, 1k reviews, 1k rebuttals and 480 meta-reviews, this is one of the largest, most complete and most diverse peer reviewing datasets to date. Learn more about the project here, or simply download the dataset and start experimenting!
How do experts reason during peer review?... This question is crucial for successful expert-AI collaboration. In the new paper, we rethink peer review as a diagnostic reasoning process. We propose Natural Language Diagnostic Abductive Reasoning as a new family of text-based reasoning tasks, where experts analyze a text step by step to arrive at a verdict. Our unique dataset of over 4000k reasoning steps opens new frontiers in the study of expert-AI collaboration. Have a look at the preprint to learn more!
Team
Dennis Zyska
PhD Student
Sheng Lu
PhD Student
Serwar Basch
PhD Student
Funding