| [1] |
ANIL R, DAI A M, FIRAT O, et al. PaLM 2 technical report[EB/OL]. 2023: arXiv: 2305.10403.
|
| [2] |
Chowdhery A, Narang S, Devlin J, et al. Palm: Scaling language modeling with pathways[J]. Journal of machine learning research, 2023, 24(240): 1-113.
|
| [3] |
OPENAI, ACHIAM J, ADLER S, et al. GPT-4 technical report[EB/OL]. 2023: arXiv: 2303.08774.
|
| [4] |
JIN D, PAN E, OUFATTOLE N, et al. What disease does this patient have a large-scale open domain question answering dataset from medical exams[J]. Applied Sciences, 2021, 11(14): 6421.
|
| [5] |
FEI Z, SHEN X, ZHU D, et al. Lawbench: Benchmarking legal knowledge of large language models[C]. //Proceedings of the 2024 conference on empirical methods in natural language processing. 2024: 7933-7962.
|
| [6] |
BANG Y, CAHYAWIJAYA S, LEE N, et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity[C]. //Proceedings of the 13th international joint conference on natural language processing and the 3rd conference of the asia-pacific chapter of the association for computational linguistics (volume 1: Long papers). 2023: 675-718.
|
| [7] |
JI Z W, LEE N, FRIESKE R, et al. Survey of hallucination in natural language generation[J]. ACM Computing Surveys, 2023, 55(12): 1-38.
|
| [8] |
MANAKUL P, LIUSIE A, GALES M. SelfCheck-Eval: A multi-module framework for zero-resource hallucination detection in large language models[J]. Patterns (N Y). 2026, 7(6): 101569.
|
| [9] |
LIN S, HILTON J, EVANS O. TruthfulQA: Measuring How Models Mimic Human Falsehoods[C]. //60th annual meeting of the Association for Computational Linguistics. Long papers, vol. 5: 60th annual meeting of the Association for Computational Linguistics (ACL 2022), Dublin, Ireland. 2022: 3214-3252.
|
| [10] |
TEAM P, DU X R, YAO Y F, et al. SuperGPQA: scaling LLM evaluation across 285 graduate disciplines[EB/OL]. 2025: arXiv: 2502.14739.
|
| [11] |
HENDRYCKS D, BURNS C, BASART S, et al. Measuring massive multitask language understanding[EB/OL]. 2020: arXiv: 2009.03300.
|
| [12] |
HUANG Y Z, BAI Y Z, ZHU Z H, et al. C-eval: a multi-level multi-discipline Chinese evaluation suite for foundation models[EB/OL]. 2023: arXiv: 2305.08322.
|
| [13] |
WANG A, SINGH A, MICHAEL J, et al. GLUE: A multi-task benchmark and analysis platform for natural language understanding[C]. //Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP. 2018: 353-355.
|
| [14] |
WANG A, PRUKSACHATKUN Y, NANGIA N, et al. Superglue: A stickier benchmark for general-purpose language understanding systems[J]. Advances in neural information processing systems, 2019, 32.
|
| [15] |
SRIVASTAVA A, RASTOGI A, RAO A, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models[J]. Transactions on machine learning research, 2023.
|
| [16] |
SUZGUN M, SCALES N, SCHÄRLI N, et al. Challenging big-bench tasks and whether chain-of-thought can solve them[C]. //Findings of the Association for Computational Linguistics: ACL 2023. 2023: 13003-13051.
|
| [17] |
CHEN M, TWOREK J, JUN H, et al. Evaluating large language models trained on code[EB/OL]. 2021: arXiv: 2107.03374.
|
| [18] |
CHUNG H W, HOU L, LONGPRE S, et al. Scaling instruction-finetuned language models[J]. Journal of Machine Learning Research, 2024, 25(70): 1-53.
|
| [19] |
OUYANG L, WU J, JIANG X, et al. Training language models to follow instructions with human feedback[J]. Advances in neural information processing systems, 2022, 35: 27730-27744.
|
| [20] |
BAI Y T, KADAVATH S, KUNDU S, et al. Constitutional AI: harmlessness from AI feedback[EB/OL]. 2022: arXiv: 2212.08073.
|
| [21] |
RAFAILOV R, SHARMA A, MITCHELL E, et al. Direct preference optimization: Your language model is secretly a reward model[J]. Advances in neural information processing systems, 2023, 36: 53728-53741.
|
| [22] |
LONGPRE S, HOU L, VU T, et al. The flan collection: Designing data and methods for effective instruction tuning[C]. //International conference on machine learning. PMLR, 2023: 22631-22648.
|
| [23] |
LEWIS P, PEREZ E, PIKTUS A, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks[J]. Advances in neural information processing systems, 2020, 33: 9459-9474.
|
| [24] |
IZACARD G, GRAVE E. Leveraging passage retrieval with generative models for open domain question answering[C]. //16th Conference of the European Chapter of the Association for Computational Linguistics, vol. 2: 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2021), Kiev, Ukraine: Association for Computational Linguistics, 2021:874-880.
|
| [25] |
KARPUKHIN V, OGUZ B, MIN S, et al. Dense passage retrieval for open-domain question answering[C]. //Conference on Empirical Methods in Natural Language Processing, vol. 11: Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), Association for Computational Linguistics, 2020: 6769-6781.
|
| [26] |
BORGEAUD S, MENSCH A, HOFFMANN J, et al. Improving language models by retrieving from trillions of tokens[C]//International conference on machine learning. PMLR, 2022: 2206-2240.
|
| [27] |
LI Y N, ZHANG W Z, YANG Y Y, et al. A survey of RAG-reasoning systems in large language models[C]. //Findings of the Association for Computational Linguistics: EMNLP 2025. Suzhou, China. Stroudsburg, PA, USA: ACL, 2025: 12120-12145.
|
| [28] |
ZHENG L, CHIANG W L, SHENG Y, et al. Judging llm-as-a-judge with mt-bench and chatbot arena[J]. Advances in neural information processing systems, 2023, 36: 46595-46623.
|
| [29] |
LIU Y, ITER D, XU Y, et al. G-eval: NLG evaluation using gpt-4 with better human alignment[C]. //Proceedings of the 2023 conference on empirical methods in natural language processing. 2023: 2511-2522.
|
| [30] |
TAN S, ZHUANG S, MONTGOMERY K, et al. Judgebench: A benchmark for evaluating llm-based judges[J]. arXiv preprint arXiv:2410.12784, 2024.
|
| [31] |
GURRAM B. Evaluating tool-using language agents: judge reliability, propagation cascades, and runtime mitigation in AgentProp-bench[EB/OL]. 2026: arXiv: 2604.16706.
|
| [32] |
GEBRU T, MORGENSTERN J, VECCHIONE B, et al. Datasheets for datasets[C]. //Communications of the ACM. ACM, 2021: 86-92.
|
| [33] |
TOUVRON H, LAVRIL T, IZACARD G, et al. LLaMA: open and efficient foundation language models[EB/OL]. 2023: arXiv: 2302.13971.
|
| [34] |
TOUVRON H, MARTIN L, STONE K, et al. Llama 2: open foundation and fine-tuned chat models[EB/OL]. 2023: arXiv: 2307.09288.
|
| [35] |
TEAM G, ANIL R, BORGEAUD S, et al. Gemini: a family of highly capable multimodal models[EB/OL]. 2023: arXiv: 2312.11805.
|
| [36] |
GRATTAFIORI A, DUBEY A, JAUHRI A, et al. The llama 3 herd of models[EB/OL]. 2024: arXiv: 2407.21783.
|
| [37] |
YANG A, LI A F, YANG B S, et al. Qwen3 technical report[EB/OL]. 2025: arXiv: 2505.09388.
|
| [38] |
GUO D Y, YANG D J, ZHANG H W, et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning[J]. Nature, 2025, 645(8081): 633-638.
|
| [39] |
WEI J, WANG X, SCHUURMANS D, et al. Chain-of-thought prompting elicits reasoning in large language models[J]. Advances in neural information processing systems, 2022, 35: 24824-24837.
|
| [40] |
WANG X, WEI J, SCHUURMANS D, et al. Self-consistency improves chain of thought reasoning in language models[J]. arXiv preprint arXiv:2203.11171, 2022.
|
| [41] |
GAO L, BIDERMAN S, BLACK S, et al. The pile: an 800GB dataset of diverse text for language modeling[EB/OL]. 2020: arXiv: 2101.00027.
|
| [42] |
CHANG J. A Risk–Utility Optimization Framework for Governing Large Language Model Responses[J]. Review of Resp AI, 2026, 1(1): 8-16.
|
| [43] |
BEAN A M, PAYNE R E, PARSONS G, et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study[J]. Nature Medicine, 2026, 32(2): 609-615.
|
| [44] |
HAPPE A, KAPLAN A, CITO J. LLMs as hackers: autonomous linux privilege escalation attacks[J]. Empirical Software Engineering, 2026, 31(3): 70.
|
| [45] |
DETTMAN J, LATHROP E, ATTAL-JUNCQUA A, et al. Prioritizing feasible and impactful actions to enable secure AI development and use in biology[J]. Biotechnology and Bioengineering, 2026: 70132.
|
| [46] |
VATSAL S, DUBEY H, SINGH A. Agentic AI in healthcare and medicine: a seven-dimensional taxonomy for empirical evaluation of LLM-based agents[J]. IEEE Access, 2026, 14: 4840-4863.
|
| [47] |
LIU X N, YANG X, LI Z K, et al. AgentHallu: benchmarking automated hallucination attribution of LLM-based agents[EB/OL]. 2026: arXiv: 2601.06818.
|
| [48] |
ROBERTSON A, LIANG H, GANI M, et al. KGHaluBench: A knowledge graph-based hallucination benchmark for evaluating the breadth and depth of llm knowledge[C]. //Findings of the Association for Computational Linguistics: EACL 2026. 2026: 3975-3989.
|
| [49] |
ABDALJALIL S, SHARMA P, SERPEDIN E, et al. Halluverse-M^3: a multitask multilingual benchmark for hallucination in LLMs[EB/OL]. 2026: arXiv: 2602.06920.
|
| [50] |
DANG Q A, NGO C, HY T S. RedBench: a universal dataset for comprehensive red teaming of large language models[EB/OL]. 2026: arXiv: 2601.03699.
|
| [51] |
PU R, LI C Z, HA R, et al. MirrorShield: towards dynamic adaptive defense against jailbreaks via entropy-guided mirror crafting[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2026, 40(39): 32746-32754.
|
| [52] |
Schick T, Dwivedi-Yu J, Dessì R, et al. Toolformer: Language models can teach yourself to use tools[J]. Advances in neural information processing systems, 2023, 36: 68539-68551.
|
| [53] |
YAO S Y, ZHAO J, YU D, et al. ReAct: synergizing reasoning and acting in language models[EB/OL]. 2022: arXiv: 2210.03629.
|
| [54] |
MIALON G, DESSÌ R, LOMELI M, et al. Augmented language models: a survey[EB/OL]. 2023: arXiv: 2302.07842.
|
| [55] |
FU C Y, CHEN P X, SHEN Y H, et al. MME: a comprehensive evaluation benchmark for multimodal large language models[EB/OL]. 2023: arXiv: 2306.13394.
|
| [56] |
ASAI A, WU Z, WANG Y, ET al. Self-rag: Learning to retrieve, generate, and critique through self-reflection[C]//The Twelfth International Conference on Learning Representations. 2023.
|