v

    Intelligent breeding

    Intelligent breeding

    Default Latest Most Read
    Please wait a minute...
    For Selected: Toggle Thumbnails
    The Chinese Livestock and Poultry Breeding    2023, 19 (7): 186-192.  
    Abstract267)      PDF(pc) (33224KB)(104)       Save
    Related Articles | Metrics | Comments0
    The Chinese Livestock and Poultry Breeding    2024, 20 (6): 5-13.  
    Abstract209)      PDF(pc) (1517KB)(223)       Save
    Related Articles | Metrics | Comments0
    The Chinese Livestock and Poultry Breeding    2025, 21 (9): 66-75.  
    Abstract56)      PDF(pc) (1243KB)(17)       Save
    Related Articles | Metrics | Comments0
    The Chinese Livestock and Poultry Breeding    2025, 21 (9): 56-65.  
    Abstract48)      PDF(pc) (2113KB)(12)       Save
    Related Articles | Metrics | Comments0
    The Chinese Livestock and Poultry Breeding    2025, 21 (10): 75-86.  
    Abstract45)      PDF(pc) (1314KB)(13)       Save
    Related Articles | Metrics | Comments0
    GBNP-2026: A benchmark for evaluating large language models' professional question-answering capabilities in animal genetics, breeding, nutrition, and production
    Chengpeng Xiao, Chenglin Li, Yanru Lei, Zhao Cai, Bofei Xie, Kejun Wang, Yujie Guo, Yujie Gong, Donghua Li, Weiwei Jin, Guirong Sun, Yan Luo, Xinjian Gu, Mingzhou Lu, Xiangtao Kang, Wenting Li
    Chinese Livestock and Poultry Breeding    2026, 22 (7): 17-27.   DOI: 10.19543/j.cnki.1673-4556.20260507.001
    Abstract31)   HTML6)    PDF(pc) (1082KB)(1)       Save

    Objective Livestock-related QA is terminology-heavy and often involves indicators, parameters, and trade-offs in breeding and production plans. This study evaluates LLM professional usability using subjective responses and explicit scoring rubrics. Methods We built a Chinese vertical benchmark, the Genetics, Breeding, Nutrition, and Production Benchmark (GBNP-2026), with 510 subjective items (363 short-answer and 147 essay) spanning the same four domains, each with a reference answer, checkable scoring points, and domain tags. Nine open- and closed-source models were tested under a unified zero-shot, context-free protocol, scored by coverage and quality, followed by error attribution on low-score cases. Results Mean scores ranged from 64.03 to 83.99 and coverage from 70.6% to 91.9%. Essay items scored higher than short-answer items on average; breeding was the weakest domain for most models. Among 628 low-score samples, missing knowledge (37.3%) and hallucinations (25.0%) dominated. Missing-knowledge rates differed across domains (χ²=8.42, P = 0.038), highest in breeding (47.1%). Conclusion Compared with multiple-choice accuracy alone, coverage-plus-quality scoring helps separate capability gaps in this setting; strictly numerical or proof-style tasks still need a dedicated sub-benchmark. GBNP-2026 and the pipeline support reproducible comparison, error diagnosis, and data planning for SFT/RAG.

    Table and Figures | Reference | Related Articles | Metrics | Comments0
    The Chinese Livestock and Poultry Breeding    2026, 22 (3): 71-81.  
    Abstract29)      PDF(pc) (996KB)(10)       Save
    Related Articles | Metrics | Comments0