Research Question
Does confidence track correctness in the same way for humans and language models? Project SOCRATES asks this question with a shared True/False task and a shared confidence scale. The goal is not only to ask which group is more accurate, but whether confidence rises and falls with correctness. Reliable AI should communicate uncertainty: an answer that sounds certain when it is wrong can be more harmful than an openly tentative one.
Why It Matters
Accuracy, calibration, and understanding are different. A system can score highly while giving nearly identical confidence to correct and incorrect answers; it can also be well calibrated in aggregate without knowing why an individual answer is right. SOCRATES therefore treats reported confidence as behavioral evidence. It does not claim that model confidence reveals an internal faculty equivalent to human metacognition.
My Role and Research Process
This was my independent research project. I developed it through a workflow of question-bank design, human questionnaires and model evaluation, aggregate analysis, and research communication. This page focuses on verified project-level findings and does not publish participant records or claim sole authorship of every technical component in the pipeline.
Study Design
The bilingual bank contained 400 True/False questions: 100 ordinary-knowledge items, 100 disciplinary items, 100 plausible-sounding pseudoscientific traps, and 100 hallucination items mixing fictional entities with real but obscure ones. After each answer, humans and models used the same five-level retrospective confidence format. Each model received four stochastic samples per question. The aggregate dataset covers 46 humans, 1,647 valid answers, and 20 model configurations. Humans answered in Chinese, while models answered in English—a central confound in any comparison.
Findings and Evidence
Human overall accuracy was 73.5%, while most model configurations scored 97–100%. Across the full 100-item hallucination category, human accuracy was 50.9%. These are first-order accuracy results, not measures of how well confidence tracked errors.
M-ratio (meta-d′/d′) measures behaviorally how well confidence separates correct from incorrect answers relative to first-order discrimination. The human estimate was 1.37, with a 95% clustered-bootstrap interval of [1.205, 1.543]. The four error-rich, data-driven model estimates ranged from 0.31 to 0.69. The figure’s legacy label “metacognitive efficiency” is behavioral shorthand; it should not be read as evidence of an internal faculty.
Within the fictional-only subset of the hallucination category, humans scored 26.9%, while model configurations scored 85–100%. Humans lowered their confidence but often affirmed fictional claims. Aggregate signal-detection analysis indicates an affirmative response bias with near-zero discrimination between fictional and real-obscure items. It does not show negative discrimination or an absence of internal metacognition.
Near-ceiling models made too few errors to support reliable M-ratio estimates. A fitted value near one in such a group cannot establish strong metacognition. The evidence tiers were at least 30 errors for data-driven estimates, 10–29 for regularized estimates, and fewer than 10 for prior-dominated estimates.
Limitations and Reflection
The human sample was small, mostly students, and drawn from one Chinese-speaking population. Language was confounded with group, and the strongest models saturated this bank. The conclusions are behavioral only. My central lesson is that accuracy, calibration, and internal understanding must not be collapsed into one idea, and that every inference should stay within what the measurement can support.
What I Can Do Next
I plan to investigate when AI systems produce unsupported answers and how stated confidence relates to those errors. I also want to test an answer-first, explanation-second format to ask whether a correct answer is supported by understanding. The explanation-scoring method still needs to be designed, and fluent language alone will not count as understanding. Finally, I aim to participate in a hackathon at Stanford, build alongside others, and meet people with shared interests; this is a future goal, not a confirmed event, admission, or affiliation.
Source note: figures and numerical claims on this page come from the project’s aggregate analysis.
研究问题
人类与语言模型的信心是否以相同方式追踪答案正确性?Project SOCRATES 用共同的判断题任务与共同的置信度量表研究这个问题。研究不仅比较谁答得更准,还考察信心能否随正确与错误而变化。可靠的 AI 应该表达不确定性:错误答案如果听起来十分确定,往往比坦率承认犹豫更危险。
为什么重要
准确率、校准和理解并不是同一件事。一个系统可能准确率很高,却对正确和错误答案给出几乎相同的信心;它也可能在总体上校准良好,却不知道某个答案为什么正确。因此,SOCRATES 将报告出来的信心视为行为证据,而不把模型信心解释为与人类元认知等同的内部能力。
项目角色与研究过程
这是我的独立研究项目。我按照题库设计→人类问卷与模型评估→聚合分析→研究传播的流程推进。本页只呈现经过核实的项目级结果,不公开参与者记录,也不声称每一项技术流程都由我独立实现。
实验设计
双语题库包含 400 道判断题:100 道常规知识题、100 道学科题、100 道表面可信的伪科学陷阱题,以及 100 道将虚构实体和真实冷僻实体混合的幻觉题。每次作答后,人类和模型都使用相同的五档回顾性信心量表;模型每题进行四次随机采样。聚合数据涵盖 46 名人类参与者、1,647 个有效答案和 20 个模型配置。人类作答中文题,模型作答英文题,这是所有直接比较中的核心混淆因素。
结果与证据
人类总体准确率为 73.5%,多数模型配置达到 97–100%。在完整的 100 道幻觉题类别中,人类准确率为 50.9%。这些是一阶准确率结果,并不衡量信心追踪错误的能力。
M-ratio(meta-d′/d′)是一项行为指标:相对于一阶判别能力,它衡量信心区分正确与错误答案的程度。人类估计值为 1.37,95% 聚类自助区间为 [1.205, 1.543];四个错误充足、由数据驱动的模型估计值为 0.31–0.69。图中沿用的“元认知效率”只是行为层面的简称,不能据此推断模型拥有某种内部能力。
在幻觉题中的纯虚构子集上,人类准确率为 26.9%,模型配置为 85–100%。人类虽然降低了信心,却仍经常肯定虚构陈述。聚合信号检测分析表明,这对应于接近零判别力下的肯定反应偏向;它既不能证明“负判别力”,也不能证明内部元认知不存在。
接近满分的模型产生的错误太少,无法支持可靠的 M-ratio 估计。这类模型中接近 1 的拟合值不能证明其具有很强的元认知。证据分层为:错误数至少 30 时属于数据驱动估计,10–29 时属于正则化估计,少于 10 时主要由先验决定。
局限与反思
人类样本较小,且主要由学生构成,来自同一个中文人群;语言与组别相互混淆;最强模型又在这套题库上达到测量天花板。所有结论都只停留在行为层面。我最重要的反思是:准确率、校准和内部理解不能被合并成同一个概念,研究推论必须受测量边界约束。
下一步探索
我计划继续研究 AI 系统何时会给出缺乏支持的答案,以及它所表达的信心如何对应这些错误。我也想探索“先回答、后解释理由”的实验形式,评估正确答案是否得到理解的支持;解释评分方法仍有待设计,表达流畅本身不能证明理解。最后,我希望未来参加在斯坦福举办的黑客松,与他人一起构建项目,并认识兴趣相近的人;这是一项目标,不代表已有确定活动、录取或隶属关系。
来源说明:本页图表与数值均来自项目的聚合分析。