AK
A.T. Kuruvilla
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
1 records found
1
Teaching Machines to Critique Computer Science Theses
A Human-in-the-Loop Framework for Discipline-Aware, Span-Anchored LLM Feedback in Thesis Supervision
To produce reviewable draft feedback on long computer-science thesis drafts under supervisor control, we present a controlled three-way comparison of LLM generation strategies on open-weight backbones, paired with a supervisor-facing PDF review interface. We compare base and LoRA-fine-tuned Llama 3.3 70B and Qwen 3.5 27B across whole-document, section-aware two-stage, and agentic review strategies, evaluated through an LLM-as-judge benchmark over all twelve configurations, a blind human rating study, and an interface usability study. Whole-document generation obtains the highest judge-macro score. Fine-tuning helps in only one of six matched base-versus-fine-tuned comparisons, indicating that supervised adaptation on a small span-anchored corpus shifts response distribution without expanding critique ability. In the blind study, the deployed section-aware generation is rated significantly higher than held-out supervisor feedback on change clarity (p = .027) and specificity (p = .018), while correctness shows no significant difference (p = .516). On the comments scored by both, three independent LLM judges rate the system above the human raters on every dimension (mean bias 0.41 points), most strongly on supervisor suitability. The PDF review interface reaches a mean System Usability Scale (SUS) score of 77.5. The results support using open-weight LLMs to produce reviewable draft feedback on long technical theses under supervisor inspection, editing, and export control.
...
To produce reviewable draft feedback on long computer-science thesis drafts under supervisor control, we present a controlled three-way comparison of LLM generation strategies on open-weight backbones, paired with a supervisor-facing PDF review interface. We compare base and LoRA-fine-tuned Llama 3.3 70B and Qwen 3.5 27B across whole-document, section-aware two-stage, and agentic review strategies, evaluated through an LLM-as-judge benchmark over all twelve configurations, a blind human rating study, and an interface usability study. Whole-document generation obtains the highest judge-macro score. Fine-tuning helps in only one of six matched base-versus-fine-tuned comparisons, indicating that supervised adaptation on a small span-anchored corpus shifts response distribution without expanding critique ability. In the blind study, the deployed section-aware generation is rated significantly higher than held-out supervisor feedback on change clarity (p = .027) and specificity (p = .018), while correctness shows no significant difference (p = .516). On the comments scored by both, three independent LLM judges rate the system above the human raters on every dimension (mean bias 0.41 points), most strongly on supervisor suitability. The PDF review interface reaches a mean System Usability Scale (SUS) score of 77.5. The results support using open-weight LLMs to produce reviewable draft feedback on long technical theses under supervisor inspection, editing, and export control.