Cem Uluoğlakçı, Benchmarking and Reducing Hallucination in Large Language Models through Hypothetical Terms
Large language models often answer confidently even when they lack relevant knowledge. This thesis asks whether they can learn a crucial form of epistemic humility: recognizing and admitting the limits of what they know. It introduces hypothetical terms, plausible concepts screened across multiple sources and likely absent from training data, to measure how readily models invent unsupported explanations. The resulting benchmarks reveal confabulation across model families, sizes, and reasoning systems. Targeted fine-tuning reduces this tendency and improves factuality across three architectures. Finally, the thesis characterizes how learned uncertainty is represented inside models and when it shapes their responses.
Date: 31.08.2026 / 14:00 Place: A-212









