中国畜禽种业 ›› 2026, Vol. 22 ›› Issue (7): 17-27.doi: 10.19543/j.cnki.1673-4556.20260507.001cstr: 32418.14.j.cnki.1673-4556.20260507.001

所属专题: 智能育种

• 前沿技术 • 上一篇    下一篇

GBNP-2026:面向畜禽遗传、育种、营养与生产的大语言模型专业问答能力评测基准

肖成朋1,2(), 李成林1,3, 雷艳茹1,2, 蔡钊1,2, 谢伯斐1,2, 王克君1,2, 郭玉洁1,2, 宫玉杰1,2, 李东华1,2, 金炜炜1,2, 孙桂荣1,2, 罗艳3, 顾兴健3, 陆明洲3, 康相涛1,2, 李文婷1,2()   

  1. 1. 河南农业大学动物科技学院,河南 郑州 450046
    2. 神农种业实验室,河南 郑州 450046
    3. 南京农业大学人工智能学院,江苏 南京 210000
  • 收稿日期:2026-02-27 出版日期:2026-07-26 发布日期:2026-07-18
  • 通讯作者: 李文婷 E-mail:18790301745@163.com;liwenting_5959@hotmail.com
  • 作者简介:

    肖成朋(1999—)男,河南驻马店人,研究方向:家禽遗传育种与繁殖,E-mail:

  • 基金资助:
    国家自然科学基金(32422081); 中原英才计划(育才系列):青年拔尖人才

GBNP-2026: A benchmark for evaluating large language models' professional question-answering capabilities in animal genetics, breeding, nutrition, and production

Chengpeng Xiao1,2(), Chenglin Li1,3, Yanru Lei1,2, Zhao Cai1,2, Bofei Xie1,2, Kejun Wang1,2, Yujie Guo1,2, Yujie Gong1,2, Donghua Li1,2, Weiwei Jin1,2, Guirong Sun1,2, Yan Luo3, Xinjian Gu3, Mingzhou Lu3, Xiangtao Kang1,2, Wenting Li1,2()   

  1. 1. College of Animal Science and Technology, Henan Agricultural University, Zhengzhou, 450046, Henan
    2. The Shennong Laboratory, Zhengzhou, 450046, Henan
    3. College of Artificial Intelligence, Nanjing Agricultural University, Nanjing, 210000, Jiangsu
  • Received:2026-02-27 Online:2026-07-26 Published:2026-07-18
  • Contact: Wenting Li E-mail:18790301745@163.com;liwenting_5959@hotmail.com

摘要:

目的 畜禽领域问答往往术语密集、推理链条长,且常涉及指标、参数与方案的权衡;本文从主观题作答与评分细则出发,考查大语言模型在畜禽领域场景中的专业可用性。 方法 构建中文垂直测评基准,即畜禽遗传、育种、营养与生产测评基准(Genetics,breeding,nutrition,and production benchmark,GBNP-2026),收录主观题510道(简答363道、论述147道);每题配有标准答案、可核查给分点与领域标签。在零样本、无上下文条件下,对9个开源与闭源模型统一评测,采用“给分点覆盖率+质量分”双维度自动评分,并对低分样本作错误归因与跨领域比较。 结果 模型总体得分呈梯度分布,平均分64.03~83.99,覆盖率70.6%~91.9%;论述题得分整体高于简答题;育种领域在多数模型中得分最低。低分样本628例中,知识遗漏占37.3%,幻觉编造占25.0%;知识遗漏比例在领域间差异显著(χ²=8.42,P=0.038)。 结论 与单纯依赖选择题准确率相比,融合覆盖率与质量分的主观题评测有助于区分模型在畜禽专业场景中的能力差异与薄弱环节,但对严格数值推导类任务仍需另行设计题型加以验证。

关键词: 大语言模型, 题库评估, 遗传育种, 自动评分, 领域评测

Abstract:

Objective Livestock-related QA is terminology-heavy and often involves indicators, parameters, and trade-offs in breeding and production plans. This study evaluates LLM professional usability using subjective responses and explicit scoring rubrics. Methods We built a Chinese vertical benchmark, the Genetics, Breeding, Nutrition, and Production Benchmark (GBNP-2026), with 510 subjective items (363 short-answer and 147 essay) spanning the same four domains, each with a reference answer, checkable scoring points, and domain tags. Nine open- and closed-source models were tested under a unified zero-shot, context-free protocol, scored by coverage and quality, followed by error attribution on low-score cases. Results Mean scores ranged from 64.03 to 83.99 and coverage from 70.6% to 91.9%. Essay items scored higher than short-answer items on average; breeding was the weakest domain for most models. Among 628 low-score samples, missing knowledge (37.3%) and hallucinations (25.0%) dominated. Missing-knowledge rates differed across domains (χ²=8.42, P = 0.038), highest in breeding (47.1%). Conclusion Compared with multiple-choice accuracy alone, coverage-plus-quality scoring helps separate capability gaps in this setting; strictly numerical or proof-style tasks still need a dedicated sub-benchmark. GBNP-2026 and the pipeline support reproducible comparison, error diagnosis, and data planning for SFT/RAG.

Key words: Large language model, Question bank assessment, Genetic breeding, Automatic scoring, Domain evaluation

中图分类号: 

  • S81

表1

主流评测基准特征对比"

基准名称Benchmark name 领域Field 题型Question type 语言Language 评测方式Evaluation method
MMLU 通用57科 选择题 英文 准确率
C-Eval 通用52科 选择题 中文 准确率
MedQA 医学 选择题 英/中 准确率
LawBench 法律 混合 中文 准确率/F1/ROUGE-L/回归距离等
AGIEval 综合考试 选择+填空 中/英 准确率/Exact Match
GBNP-2026(本文) 遗传、育种、营养、生产 简答+论述 中文 LLM评分

表2

待测模型规格一览"

模型Model 厂商Manufacturer 类型Type 上下文长度Sequence length 特点Features
GPT-5.2 OpenAI 闭源 ~128 K 综合能力最强
Kimi-2.5 Moonshot 开源 ~256 K 中文长文本优化
Gemini-3 Pro Google 闭源 ~1 M 多模态原生支持
Claude Sonnet-4.5 Anthropic 闭源 ~200 K 安全对齐优秀
GLM-4.7 智谱AI 闭源 ~128 K 中文理解优化
DeepSeek-Reasoner(R1) DeepSeek 开源 ~128 K 推理链增强
Grok-4 xAI 闭源 ~131 K 实时信息检索
Qwen-Max 阿里巴巴 开源 ~256 K 多语言支持
Qwen3-32B 阿里巴巴 开源 ~128 K 开源标杆

图1

GBNP-2026中各知识领域题型构成堆叠柱状图"

图2

GBNP-2026中单题给分点数量分布直方图"

图3

各模型总体平均得分及标准差"

图4

各模型及格率与优秀率对比"

图5

不同题型下各模型平均得分分布"

图6

知识领域模型得分热力图"

图7

各模型给分点覆盖率均值对比"

图8

低分样本错误类型分布水平柱状图(N=628)"

表3

各模型主导错误类型"

错误类型

Error type

定义

Definition

遗传Genetics

(n=168)/%

营养Nutrition

(n=187)/%

生产Production

(n=152)/%

育种Breeding

(n=121)/%

χ² PP-value
知识遗漏Missing knowledge 模型未涉及题目核心要点,知识库中缺乏相关内容 31.50 35.80 38.20 47.10 8.42 0.038
幻觉编造Hallucination 模型生成了事实性错误内容,如捏造概念或数据 29.20 23.50 25.70 21.50 3.21 0.361
答非所问Irrelevant response 模型回答了与题目无关内容,常见于术语歧义 19.60 17.60 14.50 12.40 4.15 0.246
逻辑错误Logical error 模型推理步骤或因果关系出现错误 10.10 11.80 13.80 16.50 3.87 0.275
表述模糊Ambiguous statement 模型回答不够深入,仅提到关键词未展开 7.10 8.60 5.30 2.50 6.23 0.101
其他Others 无法明确归类的错误 2.40 2.70 2.60 0.00 - -

表4

各错误类型典型案例汇编"

错误类型Error type

问题示例

Example problem

模型错误回答摘要

Summary of the model's erroneous answers

标准答案要点

Key points of the standard answer

失分原因

Reasons for point loss

知识遗漏

Missing Knowledge

简述BLUP育种值估计的基本原理 仅提到“最佳线性无偏预测”名词 Henderson混合模型方程、亲缘矩阵、固定效应与随机效应分离 缺乏数量遗传学深度知识
什么是基因组选择(GS)中的参考群体? 回答了GWAS的参考群体概念 具有表型和基因型数据的个体集合、用于训练预测模型 混淆GS与GWAS概念
列举3种常用的禽类DNA分子标记 SSR、AFLP、SNP 还应包括微卫星标记、RFLP、STR等传统标记 只答新型标记,遗漏经典标记
幻觉编造Hallucination 猪的主要经济性状遗传力范围 “产仔数遗传力0.6~0.8” 产仔数遗传力约0.1~0.15,属于低遗传力性状 严重高估,与教科书相悖
什么是杂种优势的分子机制? “杂种优势主要由显性效应导致” 显性假说、超显性假说、上位性假说,机制尚无定论 将假说当作定论
简述鸡的性别决定机制 “与哺乳动物相同,雄性为XY” 鸡为ZW型,雌性为异配ZW,雄性为同配ZZ 基础知识错误

答非所问

Irrelevant response

解释遗传漂变对小群体的影响 详细论述了自然选择的作用 等位基因频率随机波动、纯合度增加、遗传多样性丧失 完全偏题
什么是近交系数? 回答了亲缘系数的定义 个体2个等位基因来自共同祖先的概率 混淆F与R

图9

模型×错误类型分布热力图(模型内行百分比)"

[1]
ANIL R, DAI A M, FIRAT O, et al. PaLM 2 technical report[EB/OL]. 2023: arXiv: 2305.10403.
[2]
Chowdhery A, Narang S, Devlin J, et al. Palm: Scaling language modeling with pathways[J]. Journal of machine learning research, 2023, 24(240): 1-113.
[3]
OPENAI, ACHIAM J, ADLER S, et al. GPT-4 technical report[EB/OL]. 2023: arXiv: 2303.08774.
[4]
JIN D, PAN E, OUFATTOLE N, et al. What disease does this patient have a large-scale open domain question answering dataset from medical exams[J]. Applied Sciences, 2021, 11(14): 6421.
[5]
FEI Z, SHEN X, ZHU D, et al. Lawbench: Benchmarking legal knowledge of large language models[C]. //Proceedings of the 2024 conference on empirical methods in natural language processing. 2024: 7933-7962.
[6]
BANG Y, CAHYAWIJAYA S, LEE N, et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity[C]. //Proceedings of the 13th international joint conference on natural language processing and the 3rd conference of the asia-pacific chapter of the association for computational linguistics (volume 1: Long papers). 2023: 675-718.
[7]
JI Z W, LEE N, FRIESKE R, et al. Survey of hallucination in natural language generation[J]. ACM Computing Surveys, 2023, 55(12): 1-38.
[8]
MANAKUL P, LIUSIE A, GALES M. SelfCheck-Eval: A multi-module framework for zero-resource hallucination detection in large language models[J]. Patterns (N Y). 2026, 7(6): 101569.
[9]
LIN S, HILTON J, EVANS O. TruthfulQA: Measuring How Models Mimic Human Falsehoods[C]. //60th annual meeting of the Association for Computational Linguistics. Long papers, vol. 5: 60th annual meeting of the Association for Computational Linguistics (ACL 2022), Dublin, Ireland. 2022: 3214-3252.
[10]
TEAM P, DU X R, YAO Y F, et al. SuperGPQA: scaling LLM evaluation across 285 graduate disciplines[EB/OL]. 2025: arXiv: 2502.14739.
[11]
HENDRYCKS D, BURNS C, BASART S, et al. Measuring massive multitask language understanding[EB/OL]. 2020: arXiv: 2009.03300.
[12]
HUANG Y Z, BAI Y Z, ZHU Z H, et al. C-eval: a multi-level multi-discipline Chinese evaluation suite for foundation models[EB/OL]. 2023: arXiv: 2305.08322.
[13]
WANG A, SINGH A, MICHAEL J, et al. GLUE: A multi-task benchmark and analysis platform for natural language understanding[C]. //Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP. 2018: 353-355.
[14]
WANG A, PRUKSACHATKUN Y, NANGIA N, et al. Superglue: A stickier benchmark for general-purpose language understanding systems[J]. Advances in neural information processing systems, 2019, 32.
[15]
SRIVASTAVA A, RASTOGI A, RAO A, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models[J]. Transactions on machine learning research, 2023.
[16]
SUZGUN M, SCALES N, SCHÄRLI N, et al. Challenging big-bench tasks and whether chain-of-thought can solve them[C]. //Findings of the Association for Computational Linguistics: ACL 2023. 2023: 13003-13051.
[17]
CHEN M, TWOREK J, JUN H, et al. Evaluating large language models trained on code[EB/OL]. 2021: arXiv: 2107.03374.
[18]
CHUNG H W, HOU L, LONGPRE S, et al. Scaling instruction-finetuned language models[J]. Journal of Machine Learning Research, 2024, 25(70): 1-53.
[19]
OUYANG L, WU J, JIANG X, et al. Training language models to follow instructions with human feedback[J]. Advances in neural information processing systems, 2022, 35: 27730-27744.
[20]
BAI Y T, KADAVATH S, KUNDU S, et al. Constitutional AI: harmlessness from AI feedback[EB/OL]. 2022: arXiv: 2212.08073.
[21]
RAFAILOV R, SHARMA A, MITCHELL E, et al. Direct preference optimization: Your language model is secretly a reward model[J]. Advances in neural information processing systems, 2023, 36: 53728-53741.
[22]
LONGPRE S, HOU L, VU T, et al. The flan collection: Designing data and methods for effective instruction tuning[C]. //International conference on machine learning. PMLR, 2023: 22631-22648.
[23]
LEWIS P, PEREZ E, PIKTUS A, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks[J]. Advances in neural information processing systems, 2020, 33: 9459-9474.
[24]
IZACARD G, GRAVE E. Leveraging passage retrieval with generative models for open domain question answering[C]. //16th Conference of the European Chapter of the Association for Computational Linguistics, vol. 2: 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2021), Kiev, Ukraine: Association for Computational Linguistics, 2021:874-880.
[25]
KARPUKHIN V, OGUZ B, MIN S, et al. Dense passage retrieval for open-domain question answering[C]. //Conference on Empirical Methods in Natural Language Processing, vol. 11: Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), Association for Computational Linguistics, 2020: 6769-6781.
[26]
BORGEAUD S, MENSCH A, HOFFMANN J, et al. Improving language models by retrieving from trillions of tokens[C]//International conference on machine learning. PMLR, 2022: 2206-2240.
[27]
LI Y N, ZHANG W Z, YANG Y Y, et al. A survey of RAG-reasoning systems in large language models[C]. //Findings of the Association for Computational Linguistics: EMNLP 2025. Suzhou, China. Stroudsburg, PA, USA: ACL, 2025: 12120-12145.
[28]
ZHENG L, CHIANG W L, SHENG Y, et al. Judging llm-as-a-judge with mt-bench and chatbot arena[J]. Advances in neural information processing systems, 2023, 36: 46595-46623.
[29]
LIU Y, ITER D, XU Y, et al. G-eval: NLG evaluation using gpt-4 with better human alignment[C]. //Proceedings of the 2023 conference on empirical methods in natural language processing. 2023: 2511-2522.
[30]
TAN S, ZHUANG S, MONTGOMERY K, et al. Judgebench: A benchmark for evaluating llm-based judges[J]. arXiv preprint arXiv:2410.12784, 2024.
[31]
GURRAM B. Evaluating tool-using language agents: judge reliability, propagation cascades, and runtime mitigation in AgentProp-bench[EB/OL]. 2026: arXiv: 2604.16706.
[32]
GEBRU T, MORGENSTERN J, VECCHIONE B, et al. Datasheets for datasets[C]. //Communications of the ACM. ACM, 2021: 86-92.
[33]
TOUVRON H, LAVRIL T, IZACARD G, et al. LLaMA: open and efficient foundation language models[EB/OL]. 2023: arXiv: 2302.13971.
[34]
TOUVRON H, MARTIN L, STONE K, et al. Llama 2: open foundation and fine-tuned chat models[EB/OL]. 2023: arXiv: 2307.09288.
[35]
TEAM G, ANIL R, BORGEAUD S, et al. Gemini: a family of highly capable multimodal models[EB/OL]. 2023: arXiv: 2312.11805.
[36]
GRATTAFIORI A, DUBEY A, JAUHRI A, et al. The llama 3 herd of models[EB/OL]. 2024: arXiv: 2407.21783.
[37]
YANG A, LI A F, YANG B S, et al. Qwen3 technical report[EB/OL]. 2025: arXiv: 2505.09388.
[38]
GUO D Y, YANG D J, ZHANG H W, et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning[J]. Nature, 2025, 645(8081): 633-638.
[39]
WEI J, WANG X, SCHUURMANS D, et al. Chain-of-thought prompting elicits reasoning in large language models[J]. Advances in neural information processing systems, 2022, 35: 24824-24837.
[40]
WANG X, WEI J, SCHUURMANS D, et al. Self-consistency improves chain of thought reasoning in language models[J]. arXiv preprint arXiv:2203.11171, 2022.
[41]
GAO L, BIDERMAN S, BLACK S, et al. The pile: an 800GB dataset of diverse text for language modeling[EB/OL]. 2020: arXiv: 2101.00027.
[42]
CHANG J. A Risk–Utility Optimization Framework for Governing Large Language Model Responses[J]. Review of Resp AI, 2026, 1(1): 8-16.
[43]
BEAN A M, PAYNE R E, PARSONS G, et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study[J]. Nature Medicine, 2026, 32(2): 609-615.
[44]
HAPPE A, KAPLAN A, CITO J. LLMs as hackers: autonomous linux privilege escalation attacks[J]. Empirical Software Engineering, 2026, 31(3): 70.
[45]
DETTMAN J, LATHROP E, ATTAL-JUNCQUA A, et al. Prioritizing feasible and impactful actions to enable secure AI development and use in biology[J]. Biotechnology and Bioengineering, 2026: 70132.
[46]
VATSAL S, DUBEY H, SINGH A. Agentic AI in healthcare and medicine: a seven-dimensional taxonomy for empirical evaluation of LLM-based agents[J]. IEEE Access, 2026, 14: 4840-4863.
[47]
LIU X N, YANG X, LI Z K, et al. AgentHallu: benchmarking automated hallucination attribution of LLM-based agents[EB/OL]. 2026: arXiv: 2601.06818.
[48]
ROBERTSON A, LIANG H, GANI M, et al. KGHaluBench: A knowledge graph-based hallucination benchmark for evaluating the breadth and depth of llm knowledge[C]. //Findings of the Association for Computational Linguistics: EACL 2026. 2026: 3975-3989.
[49]
ABDALJALIL S, SHARMA P, SERPEDIN E, et al. Halluverse-M^3: a multitask multilingual benchmark for hallucination in LLMs[EB/OL]. 2026: arXiv: 2602.06920.
[50]
DANG Q A, NGO C, HY T S. RedBench: a universal dataset for comprehensive red teaming of large language models[EB/OL]. 2026: arXiv: 2601.03699.
[51]
PU R, LI C Z, HA R, et al. MirrorShield: towards dynamic adaptive defense against jailbreaks via entropy-guided mirror crafting[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2026, 40(39): 32746-32754.
[52]
Schick T, Dwivedi-Yu J, Dessì R, et al. Toolformer: Language models can teach yourself to use tools[J]. Advances in neural information processing systems, 2023, 36: 68539-68551.
[53]
YAO S Y, ZHAO J, YU D, et al. ReAct: synergizing reasoning and acting in language models[EB/OL]. 2022: arXiv: 2210.03629.
[54]
MIALON G, DESSÌ R, LOMELI M, et al. Augmented language models: a survey[EB/OL]. 2023: arXiv: 2302.07842.
[55]
FU C Y, CHEN P X, SHEN Y H, et al. MME: a comprehensive evaluation benchmark for multimodal large language models[EB/OL]. 2023: arXiv: 2306.13394.
[56]
ASAI A, WU Z, WANG Y, ET al. Self-rag: Learning to retrieve, generate, and critique through self-reflection[C]//The Twelfth International Conference on Learning Representations. 2023.
[1] 范津玮, 钟梓奇, 潘德优, 苏之青, 刘思宇, 杨政, 侯冠彧, 肖倩. 多组学技术在家禽遗传育种中的应用进展[J]. 中国畜禽种业, 2026, 22(7): 45-52.
[2] 程瑞琦, 周华倩, 杨华, 杨永林, 余乾, 张文喆, 陈岩, 赵宗胜, 崔蕾, 马春萍. SNP芯片开发及其在畜禽遗传育种中的应用研究进展[J]. 中国畜禽种业, 2026, 22(5): 61-69.
[3] 王添添, 卲嘉豪, 段文淼, 鲁佳宁, 曾检华, 刘小红, 袁晓龙. 毛色遗传多样性在家畜遗传育种中的研究进展[J]. 中国畜禽种业, 2026, 22(5): 108-115.
[4] 王利刚,甄霆,申峻松. 新时代背景下 《动物遗传育种》 课程思政教学改革与实践[J]. 中国畜禽种业, 2023, 19(9): 168-172.
[5] 张格阳, 吕世杰, 朱肖亭, 翟亚莹, 张志浩, 施巧婷, 张子敬, 屈俊峰, 陈付英, 王二耀. 液相芯片技术在动物遗传育种中的应用[J]. 中国畜禽种业, 2023, 19(6): 24-28.
[6] 伍昌华, 梁燕, 郑汝青. 贵州关岭黄牛种群特性研究进展及育种展望[J]. 中国畜禽种业, 2023, 19(6): 10-14.
[7] 马云, 谷帅锋, 潘翠丽, 杨梦丽, 冯兰. 宁夏回族自治区肉牛种业现状、问题及对策建议[J]. 中国畜禽种业, 2023, 19(2): 9-13.
[8] 夏·巴音克西克. 肉牛遗传育种与繁殖技术发展趋势探讨[J]. 中国畜禽种业, 2022, 18(4): 81-81.
[9] 邢杰. 随机回归模型在畜禽育种中的应用[J]. 中国畜禽种业, 2021, 17(8): 40-43.
[10] 谭萍. 肉牛遗传育种与繁殖技术发展趋势[J]. 中国畜禽种业, 2021, 17(3): 101-102.
[11] 金森. 代谢组学在畜禽遗传育种中的应用分析[J]. 中国畜禽种业, 2021, 17(10): 30-31.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!