Independent
All model APIs are independently purchased. We do not accept vendor-sponsored evaluations or PR collaborations.
Dianming AI Lab
Why can you trust the benchmark data you see? Because we don't accept vendor payments, don't modify scores, and don't hide failure cases. Every metric is reproducible, and every test set is publicly available.
OUR PRINCIPLES
We constrain our own evaluation behavior first, then discuss the strengths and weaknesses of any model.
All model APIs are independently purchased. We do not accept vendor-sponsored evaluations or PR collaborations.
Every model uses the same parameters, same test sets, and same scoring standards.
Raw data, failure cases, error rates, and confidence intervals are publicly available, supporting result reproduction.
EVALUATION MATRIX
Every model is scored using the same set of dimensions, ensuring horizontal comparability and vertical traceability.
10 / 10Capability, Performance, Cost, and Security All Included
C-Eval + Custom Hard Sentence Set
HumanEval + MBPP + Business Dataset
GSM8K + MATH + Logic Traps
NiAH 200K + Long Text Summarization
Function Calling Accuracy
Image / OCR / PDF / Charts
Thinking Mode Observability
TTFT + TPS + Decay Curve
Single Task Cost + Cache Discount
Jailbreak Resistance + Bias Detection
Multi-task language understanding covering 57 subjects. MMLU-Pro includes harder reasoning questions to better distinguish top models.
Contains 164 Python programming problems. pass@1 is the most frequently cited code capability metric in the industry.
Contains 8.5K elementary math word problems. High requirements for reasoning chain completeness, making it an important standard for testing model reasoning ability.
Covers 52 subjects including basic science, engineering, humanities, and China-specific topics, used to evaluate the model's comprehensive Chinese ability.
Inserts key information at random positions in 200K tokens of text to test whether the model can accurately find and use it.
STANDARD PROCESS
From test question preparation to publication, every step leaves auditable and traceable evaluation records.
Test questions are strictly confidential, with random tokens added to reduce the impact of training set contamination on results.
All models use the same temperature (temp=0.0), system prompts, and call parameters.
Each question is independently called 3 times. Mean and standard deviation are calculated to eliminate random deviations.
Subjective questions are scored independently by 2 annotators. Disagreements are arbitrated by a third person. Kappa ≥ 0.8.
Full raw data is publicly released, including failure cases, error rates, and confidence intervals.
AI LAB TEAM
5 full-time evaluation engineers and 2 visiting academic advisors, with professional backgrounds covering NLP, recommendation systems, financial risk control, and computational linguistics.
Lab Director · Former BAT NLP Algorithm Expert · PhD in Computational Linguistics
Evaluation System Architect · Large-scale Distributed Stress Testing · 8 Years ML Infrastructure Experience
Chinese Corpus Expert · Former Institute of Linguistics, CASS · C-Eval Co-builder
Multimodal Direction · Vision / OCR / Chart Understanding Evaluation
OUR BOUNDARIES
The clearer the boundaries, the more trustworthy the evaluation conclusions.
Evaluation is an attitude. If scores change due to business relationships, they are not scores—they are advertisements. What we produce are scores.
— Dianming AI Lab
CHOOSE WITH EVIDENCE
Based on unified standards of continuous real testing, providing more honest model selection criteria for different business scenarios.