v

Chinese Livestock and Poultry Breeding ›› 2026, Vol. 22 ›› Issue (7): 17-27.doi: 10.19543/j.cnki.1673-4556.20260507.001cstr: 32418.14.j.cnki.1673-4556.20260507.001

Special Issue: Intelligent breeding

• Advanced Technology • Previous Articles     Next Articles

GBNP-2026: A benchmark for evaluating large language models' professional question-answering capabilities in animal genetics, breeding, nutrition, and production

Chengpeng Xiao1,2(), Chenglin Li1,3, Yanru Lei1,2, Zhao Cai1,2, Bofei Xie1,2, Kejun Wang1,2, Yujie Guo1,2, Yujie Gong1,2, Donghua Li1,2, Weiwei Jin1,2, Guirong Sun1,2, Yan Luo3, Xinjian Gu3, Mingzhou Lu3, Xiangtao Kang1,2, Wenting Li1,2()   

  1. 1. College of Animal Science and Technology, Henan Agricultural University, Zhengzhou, 450046, Henan
    2. The Shennong Laboratory, Zhengzhou, 450046, Henan
    3. College of Artificial Intelligence, Nanjing Agricultural University, Nanjing, 210000, Jiangsu
  • Received:2026-02-27 Online:2026-07-26 Published:2026-07-18
  • Contact: Wenting Li E-mail:18790301745@163.com;liwenting_5959@hotmail.com

Abstract:

Objective Livestock-related QA is terminology-heavy and often involves indicators, parameters, and trade-offs in breeding and production plans. This study evaluates LLM professional usability using subjective responses and explicit scoring rubrics. Methods We built a Chinese vertical benchmark, the Genetics, Breeding, Nutrition, and Production Benchmark (GBNP-2026), with 510 subjective items (363 short-answer and 147 essay) spanning the same four domains, each with a reference answer, checkable scoring points, and domain tags. Nine open- and closed-source models were tested under a unified zero-shot, context-free protocol, scored by coverage and quality, followed by error attribution on low-score cases. Results Mean scores ranged from 64.03 to 83.99 and coverage from 70.6% to 91.9%. Essay items scored higher than short-answer items on average; breeding was the weakest domain for most models. Among 628 low-score samples, missing knowledge (37.3%) and hallucinations (25.0%) dominated. Missing-knowledge rates differed across domains (χ²=8.42, P = 0.038), highest in breeding (47.1%). Conclusion Compared with multiple-choice accuracy alone, coverage-plus-quality scoring helps separate capability gaps in this setting; strictly numerical or proof-style tasks still need a dedicated sub-benchmark. GBNP-2026 and the pipeline support reproducible comparison, error diagnosis, and data planning for SFT/RAG.

Key words: Large language model, Question bank assessment, Genetic breeding, Automatic scoring, Domain evaluation

CLC Number: 

  • S81

Table 1

Comparison of characteristics of mainstream evaluation benchmarks"

基准名称Benchmark name 领域Field 题型Question type 语言Language 评测方式Evaluation method
MMLU 通用57科 选择题 英文 准确率
C-Eval 通用52科 选择题 中文 准确率
MedQA 医学 选择题 英/中 准确率
LawBench 法律 混合 中文 准确率/F1/ROUGE-L/回归距离等
AGIEval 综合考试 选择+填空 中/英 准确率/Exact Match
GBNP-2026(本文) 遗传、育种、营养、生产 简答+论述 中文 LLM评分

Table 2

Specification overview of the test model"

模型Model 厂商Manufacturer 类型Type 上下文长度Sequence length 特点Features
GPT-5.2 OpenAI 闭源 ~128 K 综合能力最强
Kimi-2.5 Moonshot 开源 ~256 K 中文长文本优化
Gemini-3 Pro Google 闭源 ~1 M 多模态原生支持
Claude Sonnet-4.5 Anthropic 闭源 ~200 K 安全对齐优秀
GLM-4.7 智谱AI 闭源 ~128 K 中文理解优化
DeepSeek-Reasoner(R1) DeepSeek 开源 ~128 K 推理链增强
Grok-4 xAI 闭源 ~131 K 实时信息检索
Qwen-Max 阿里巴巴 开源 ~256 K 多语言支持
Qwen3-32B 阿里巴巴 开源 ~128 K 开源标杆

Fig. 1

Question type composition and scoring-point distribution of the GBNP-2026 benchmark"

Fig. 2

Histogram showing the distribution of the number of scoring points for each question in GBNP-2026"

Fig. 3

Shows the overall average scores and standard deviations of each model"

Fig. 4

Presents a comparison of the pass rates and excellent rates of each model"

Fig. 5

Distribution of average scores for each model under different question types"

Fig. 6

Heatmap of knowledge domain model scores"

Fig. 7

Comparison of the average coverage rate of scoring points for each model"

Fig. 8

Horizontal bar chart showing the distribution of error types for low-scoring samples(N=628)"

Table 3

Dominant error types of each model"

错误类型

Error type

定义

Definition

遗传Genetics

(n=168)/%

营养Nutrition

(n=187)/%

生产Production

(n=152)/%

育种Breeding

(n=121)/%

χ² PP-value
知识遗漏Missing knowledge 模型未涉及题目核心要点,知识库中缺乏相关内容 31.50 35.80 38.20 47.10 8.42 0.038
幻觉编造Hallucination 模型生成了事实性错误内容,如捏造概念或数据 29.20 23.50 25.70 21.50 3.21 0.361
答非所问Irrelevant response 模型回答了与题目无关内容,常见于术语歧义 19.60 17.60 14.50 12.40 4.15 0.246
逻辑错误Logical error 模型推理步骤或因果关系出现错误 10.10 11.80 13.80 16.50 3.87 0.275
表述模糊Ambiguous statement 模型回答不够深入,仅提到关键词未展开 7.10 8.60 5.30 2.50 6.23 0.101
其他Others 无法明确归类的错误 2.40 2.70 2.60 0.00 - -

Table 4

Compilation of typical cases of each error type"

错误类型Error type

问题示例

Example problem

模型错误回答摘要

Summary of the model's erroneous answers

标准答案要点

Key points of the standard answer

失分原因

Reasons for point loss

知识遗漏

Missing Knowledge

简述BLUP育种值估计的基本原理 仅提到“最佳线性无偏预测”名词 Henderson混合模型方程、亲缘矩阵、固定效应与随机效应分离 缺乏数量遗传学深度知识
什么是基因组选择(GS)中的参考群体? 回答了GWAS的参考群体概念 具有表型和基因型数据的个体集合、用于训练预测模型 混淆GS与GWAS概念
列举3种常用的禽类DNA分子标记 SSR、AFLP、SNP 还应包括微卫星标记、RFLP、STR等传统标记 只答新型标记,遗漏经典标记
幻觉编造Hallucination 猪的主要经济性状遗传力范围 “产仔数遗传力0.6~0.8” 产仔数遗传力约0.1~0.15,属于低遗传力性状 严重高估,与教科书相悖
什么是杂种优势的分子机制? “杂种优势主要由显性效应导致” 显性假说、超显性假说、上位性假说,机制尚无定论 将假说当作定论
简述鸡的性别决定机制 “与哺乳动物相同,雄性为XY” 鸡为ZW型,雌性为异配ZW,雄性为同配ZZ 基础知识错误

答非所问

Irrelevant response

解释遗传漂变对小群体的影响 详细论述了自然选择的作用 等位基因频率随机波动、纯合度增加、遗传多样性丧失 完全偏题
什么是近交系数? 回答了亲缘系数的定义 个体2个等位基因来自共同祖先的概率 混淆F与R

Fig. 9

Heat map of model×error type distribution(percentage in rows within the model)"

[1]
ANIL R, DAI A M, FIRAT O, et al. PaLM 2 technical report[EB/OL]. 2023: arXiv: 2305.10403.
[2]
Chowdhery A, Narang S, Devlin J, et al. Palm: Scaling language modeling with pathways[J]. Journal of machine learning research, 2023, 24(240): 1-113.
[3]
OPENAI, ACHIAM J, ADLER S, et al. GPT-4 technical report[EB/OL]. 2023: arXiv: 2303.08774.
[4]
JIN D, PAN E, OUFATTOLE N, et al. What disease does this patient have a large-scale open domain question answering dataset from medical exams[J]. Applied Sciences, 2021, 11(14): 6421.
[5]
FEI Z, SHEN X, ZHU D, et al. Lawbench: Benchmarking legal knowledge of large language models[C]. //Proceedings of the 2024 conference on empirical methods in natural language processing. 2024: 7933-7962.
[6]
BANG Y, CAHYAWIJAYA S, LEE N, et al. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity[C]. //Proceedings of the 13th international joint conference on natural language processing and the 3rd conference of the asia-pacific chapter of the association for computational linguistics (volume 1: Long papers). 2023: 675-718.
[7]
JI Z W, LEE N, FRIESKE R, et al. Survey of hallucination in natural language generation[J]. ACM Computing Surveys, 2023, 55(12): 1-38.
[8]
MANAKUL P, LIUSIE A, GALES M. SelfCheck-Eval: A multi-module framework for zero-resource hallucination detection in large language models[J]. Patterns (N Y). 2026, 7(6): 101569.
[9]
LIN S, HILTON J, EVANS O. TruthfulQA: Measuring How Models Mimic Human Falsehoods[C]. //60th annual meeting of the Association for Computational Linguistics. Long papers, vol. 5: 60th annual meeting of the Association for Computational Linguistics (ACL 2022), Dublin, Ireland. 2022: 3214-3252.
[10]
TEAM P, DU X R, YAO Y F, et al. SuperGPQA: scaling LLM evaluation across 285 graduate disciplines[EB/OL]. 2025: arXiv: 2502.14739.
[11]
HENDRYCKS D, BURNS C, BASART S, et al. Measuring massive multitask language understanding[EB/OL]. 2020: arXiv: 2009.03300.
[12]
HUANG Y Z, BAI Y Z, ZHU Z H, et al. C-eval: a multi-level multi-discipline Chinese evaluation suite for foundation models[EB/OL]. 2023: arXiv: 2305.08322.
[13]
WANG A, SINGH A, MICHAEL J, et al. GLUE: A multi-task benchmark and analysis platform for natural language understanding[C]. //Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP. 2018: 353-355.
[14]
WANG A, PRUKSACHATKUN Y, NANGIA N, et al. Superglue: A stickier benchmark for general-purpose language understanding systems[J]. Advances in neural information processing systems, 2019, 32.
[15]
SRIVASTAVA A, RASTOGI A, RAO A, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models[J]. Transactions on machine learning research, 2023.
[16]
SUZGUN M, SCALES N, SCHÄRLI N, et al. Challenging big-bench tasks and whether chain-of-thought can solve them[C]. //Findings of the Association for Computational Linguistics: ACL 2023. 2023: 13003-13051.
[17]
CHEN M, TWOREK J, JUN H, et al. Evaluating large language models trained on code[EB/OL]. 2021: arXiv: 2107.03374.
[18]
CHUNG H W, HOU L, LONGPRE S, et al. Scaling instruction-finetuned language models[J]. Journal of Machine Learning Research, 2024, 25(70): 1-53.
[19]
OUYANG L, WU J, JIANG X, et al. Training language models to follow instructions with human feedback[J]. Advances in neural information processing systems, 2022, 35: 27730-27744.
[20]
BAI Y T, KADAVATH S, KUNDU S, et al. Constitutional AI: harmlessness from AI feedback[EB/OL]. 2022: arXiv: 2212.08073.
[21]
RAFAILOV R, SHARMA A, MITCHELL E, et al. Direct preference optimization: Your language model is secretly a reward model[J]. Advances in neural information processing systems, 2023, 36: 53728-53741.
[22]
LONGPRE S, HOU L, VU T, et al. The flan collection: Designing data and methods for effective instruction tuning[C]. //International conference on machine learning. PMLR, 2023: 22631-22648.
[23]
LEWIS P, PEREZ E, PIKTUS A, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks[J]. Advances in neural information processing systems, 2020, 33: 9459-9474.
[24]
IZACARD G, GRAVE E. Leveraging passage retrieval with generative models for open domain question answering[C]. //16th Conference of the European Chapter of the Association for Computational Linguistics, vol. 2: 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2021), Kiev, Ukraine: Association for Computational Linguistics, 2021:874-880.
[25]
KARPUKHIN V, OGUZ B, MIN S, et al. Dense passage retrieval for open-domain question answering[C]. //Conference on Empirical Methods in Natural Language Processing, vol. 11: Conference on Empirical Methods in Natural Language Processing (EMNLP 2020), Association for Computational Linguistics, 2020: 6769-6781.
[26]
BORGEAUD S, MENSCH A, HOFFMANN J, et al. Improving language models by retrieving from trillions of tokens[C]//International conference on machine learning. PMLR, 2022: 2206-2240.
[27]
LI Y N, ZHANG W Z, YANG Y Y, et al. A survey of RAG-reasoning systems in large language models[C]. //Findings of the Association for Computational Linguistics: EMNLP 2025. Suzhou, China. Stroudsburg, PA, USA: ACL, 2025: 12120-12145.
[28]
ZHENG L, CHIANG W L, SHENG Y, et al. Judging llm-as-a-judge with mt-bench and chatbot arena[J]. Advances in neural information processing systems, 2023, 36: 46595-46623.
[29]
LIU Y, ITER D, XU Y, et al. G-eval: NLG evaluation using gpt-4 with better human alignment[C]. //Proceedings of the 2023 conference on empirical methods in natural language processing. 2023: 2511-2522.
[30]
TAN S, ZHUANG S, MONTGOMERY K, et al. Judgebench: A benchmark for evaluating llm-based judges[J]. arXiv preprint arXiv:2410.12784, 2024.
[31]
GURRAM B. Evaluating tool-using language agents: judge reliability, propagation cascades, and runtime mitigation in AgentProp-bench[EB/OL]. 2026: arXiv: 2604.16706.
[32]
GEBRU T, MORGENSTERN J, VECCHIONE B, et al. Datasheets for datasets[C]. //Communications of the ACM. ACM, 2021: 86-92.
[33]
TOUVRON H, LAVRIL T, IZACARD G, et al. LLaMA: open and efficient foundation language models[EB/OL]. 2023: arXiv: 2302.13971.
[34]
TOUVRON H, MARTIN L, STONE K, et al. Llama 2: open foundation and fine-tuned chat models[EB/OL]. 2023: arXiv: 2307.09288.
[35]
TEAM G, ANIL R, BORGEAUD S, et al. Gemini: a family of highly capable multimodal models[EB/OL]. 2023: arXiv: 2312.11805.
[36]
GRATTAFIORI A, DUBEY A, JAUHRI A, et al. The llama 3 herd of models[EB/OL]. 2024: arXiv: 2407.21783.
[37]
YANG A, LI A F, YANG B S, et al. Qwen3 technical report[EB/OL]. 2025: arXiv: 2505.09388.
[38]
GUO D Y, YANG D J, ZHANG H W, et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning[J]. Nature, 2025, 645(8081): 633-638.
[39]
WEI J, WANG X, SCHUURMANS D, et al. Chain-of-thought prompting elicits reasoning in large language models[J]. Advances in neural information processing systems, 2022, 35: 24824-24837.
[40]
WANG X, WEI J, SCHUURMANS D, et al. Self-consistency improves chain of thought reasoning in language models[J]. arXiv preprint arXiv:2203.11171, 2022.
[41]
GAO L, BIDERMAN S, BLACK S, et al. The pile: an 800GB dataset of diverse text for language modeling[EB/OL]. 2020: arXiv: 2101.00027.
[42]
CHANG J. A Risk–Utility Optimization Framework for Governing Large Language Model Responses[J]. Review of Resp AI, 2026, 1(1): 8-16.
[43]
BEAN A M, PAYNE R E, PARSONS G, et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study[J]. Nature Medicine, 2026, 32(2): 609-615.
[44]
HAPPE A, KAPLAN A, CITO J. LLMs as hackers: autonomous linux privilege escalation attacks[J]. Empirical Software Engineering, 2026, 31(3): 70.
[45]
DETTMAN J, LATHROP E, ATTAL-JUNCQUA A, et al. Prioritizing feasible and impactful actions to enable secure AI development and use in biology[J]. Biotechnology and Bioengineering, 2026: 70132.
[46]
VATSAL S, DUBEY H, SINGH A. Agentic AI in healthcare and medicine: a seven-dimensional taxonomy for empirical evaluation of LLM-based agents[J]. IEEE Access, 2026, 14: 4840-4863.
[47]
LIU X N, YANG X, LI Z K, et al. AgentHallu: benchmarking automated hallucination attribution of LLM-based agents[EB/OL]. 2026: arXiv: 2601.06818.
[48]
ROBERTSON A, LIANG H, GANI M, et al. KGHaluBench: A knowledge graph-based hallucination benchmark for evaluating the breadth and depth of llm knowledge[C]. //Findings of the Association for Computational Linguistics: EACL 2026. 2026: 3975-3989.
[49]
ABDALJALIL S, SHARMA P, SERPEDIN E, et al. Halluverse-M^3: a multitask multilingual benchmark for hallucination in LLMs[EB/OL]. 2026: arXiv: 2602.06920.
[50]
DANG Q A, NGO C, HY T S. RedBench: a universal dataset for comprehensive red teaming of large language models[EB/OL]. 2026: arXiv: 2601.03699.
[51]
PU R, LI C Z, HA R, et al. MirrorShield: towards dynamic adaptive defense against jailbreaks via entropy-guided mirror crafting[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2026, 40(39): 32746-32754.
[52]
Schick T, Dwivedi-Yu J, Dessì R, et al. Toolformer: Language models can teach yourself to use tools[J]. Advances in neural information processing systems, 2023, 36: 68539-68551.
[53]
YAO S Y, ZHAO J, YU D, et al. ReAct: synergizing reasoning and acting in language models[EB/OL]. 2022: arXiv: 2210.03629.
[54]
MIALON G, DESSÌ R, LOMELI M, et al. Augmented language models: a survey[EB/OL]. 2023: arXiv: 2302.07842.
[55]
FU C Y, CHEN P X, SHEN Y H, et al. MME: a comprehensive evaluation benchmark for multimodal large language models[EB/OL]. 2023: arXiv: 2306.13394.
[56]
ASAI A, WU Z, WANG Y, ET al. Self-rag: Learning to retrieve, generate, and critique through self-reflection[C]//The Twelfth International Conference on Learning Representations. 2023.
[1] Jinwei Fan, Ziqi Zhong, Deyou Pan, Zhiqing Su, Siyu Liu, Zheng Yang, Guanyu Hou, Qian Xiao. Research progress of multi-omics technology in poultry genetic breeding [J]. Chinese Livestock and Poultry Breeding, 2026, 22(7): 45-52.
[2] Tiantian Wang, Jiahao Shao, Wenmiao Duan, Jianing Lu, Jianhua Zeng, Xiaohong Liu, Xiaolong Yuan. Research progress on genetic diversity of coat color in livestock genetic breeding [J]. Chinese Livestock and Poultry Breeding, 2026, 22(5): 108-115.
[3] Zhang Geyang, Lv Shijie, Zhu Xiaoting, Zhai Yaying, Zhang Zhihao, Shi Qiaoting, Zhang Zijing, Qu Junfeng, Chen Fuying, Wang Eryao. Application of Liquid Chip Technology in Animal Genetics and Breeding [J]. The Chinese Livestock and Poultry Breeding, 2023, 19(6): 24-28.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!