Objective Livestock-related QA is terminology-heavy and often involves indicators, parameters, and trade-offs in breeding and production plans. This study evaluates LLM professional usability using subjective responses and explicit scoring rubrics. Methods We built a Chinese vertical benchmark, the Genetics, Breeding, Nutrition, and Production Benchmark (GBNP-2026), with 510 subjective items (363 short-answer and 147 essay) spanning the same four domains, each with a reference answer, checkable scoring points, and domain tags. Nine open- and closed-source models were tested under a unified zero-shot, context-free protocol, scored by coverage and quality, followed by error attribution on low-score cases. Results Mean scores ranged from 64.03 to 83.99 and coverage from 70.6% to 91.9%. Essay items scored higher than short-answer items on average; breeding was the weakest domain for most models. Among 628 low-score samples, missing knowledge (37.3%) and hallucinations (25.0%) dominated. Missing-knowledge rates differed across domains (χ²=8.42, P = 0.038), highest in breeding (47.1%). Conclusion Compared with multiple-choice accuracy alone, coverage-plus-quality scoring helps separate capability gaps in this setting; strictly numerical or proof-style tasks still need a dedicated sub-benchmark. GBNP-2026 and the pipeline support reproducible comparison, error diagnosis, and data planning for SFT/RAG.