KHABIBILLO OS
Evidence Position Sensitivity Across Document-Count Conditions in Multi-Document QA
Undergraduate Researcher · Data and Language Intelligence Lab, Kyungpook National University · Jun 2026 – Sep 2026
- Reproduced and evaluated Search-R1 search-augmented reasoning on NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, MuSiQue and Bamboogle.
- Evaluated retrieval-based reasoning (SeaKR) and tool-augmented LLMs (ToolBench / ToolLLM): tool selection, retrieval, and how reliably tool-using models produce correct results.
- Built the experiment and evaluation pipelines in Python with PyTorch, Hugging Face and vLLM, including statistical significance testing.
Evidence Position Sensitivity Across Document-Count Conditions in Multi-Document QA
Accepted for publication. Does the position of supporting evidence among retrieved documents change answer quality, and does that effect depend on how many documents there are?
Evidence placed at the beginning, middle or end of 8, 10, 15 or 20 documents, with total input length held fixed.
Read the research page