Dianming AI Lab

Independent · Neutral · Transparent
Model Evaluation Methodology

Why can you trust the benchmark data you see? Because we don't accept vendor payments, don't modify scores, and don't hide failure cases. Every metric is reproducible, and every test set is publicly available.

  • Same-parameter Horizontal Comparison
  • Double-Blind Human Annotation
  • Raw Data Reproducible
10 Evaluation Dimensions
5 Public Benchmarks
30+ Models Under Observation
Weekly Fixed Retesting

OUR PRINCIPLES

Trusted Evaluation Starts with Methodology

We constrain our own evaluation behavior first, then discuss the strengths and weaknesses of any model.

01

Independent

All model APIs are independently purchased. We do not accept vendor-sponsored evaluations or PR collaborations.

02

Neutral

Every model uses the same parameters, same test sets, and same scoring standards.

03

Transparent

Raw data, failure cases, error rates, and confidence intervals are publicly available, supporting result reproduction.

EVALUATION MATRIX

10 Evaluation Dimensions

Every model is scored using the same set of dimensions, ensuring horizontal comparability and vertical traceability.

10 / 10Capability, Performance, Cost, and Security All Included

01

Chinese Understanding

C-Eval + Custom Hard Sentence Set

02

Code Capability

HumanEval + MBPP + Business Dataset

03

Reasoning Ability

GSM8K + MATH + Logic Traps

04

Long Context

NiAH 200K + Long Text Summarization

05

Tool Calling

Function Calling Accuracy

06

Multimodal

Image / OCR / PDF / Charts

07

Reasoning Transparency

Thinking Mode Observability

08

Response Speed

TTFT + TPS + Decay Curve

09

Cost Performance

Single Task Cost + Cache Discount

10

Safety Alignment

Jailbreak Resistance + Bias Detection

01

MMLU / MMLU-Pro

Comprehensive Knowledge

Multi-task language understanding covering 57 subjects. MMLU-Pro includes harder reasoning questions to better distinguish top models.

02

HumanEval

Code Generation

Contains 164 Python programming problems. pass@1 is the most frequently cited code capability metric in the industry.

03

GSM8K

Math Reasoning

Contains 8.5K elementary math word problems. High requirements for reasoning chain completeness, making it an important standard for testing model reasoning ability.

04

C-Eval

Chinese Comprehensive

Covers 52 subjects including basic science, engineering, humanities, and China-specific topics, used to evaluate the model's comprehensive Chinese ability.

05

NiAH 200K

Long Context Recall

Inserts key information at random positions in 200K tokens of text to test whether the model can accurately find and use it.

STANDARD PROCESS

5-Step Standard Testing Process

From test question preparation to publication, every step leaves auditable and traceable evaluation records.

  1. 01

    Data Preparation

    Test questions are strictly confidential, with random tokens added to reduce the impact of training set contamination on results.

  2. 02

    Parameter Alignment

    All models use the same temperature (temp=0.0), system prompts, and call parameters.

  3. 03

    Three-Round Sampling

    Each question is independently called 3 times. Mean and standard deviation are calculated to eliminate random deviations.

  4. 04

    Double-Blind Scoring

    Subjective questions are scored independently by 2 annotators. Disagreements are arbitrated by a third person. Kappa ≥ 0.8.

  5. 05

    Publish Data

    Full raw data is publicly released, including failure cases, error rates, and confidence intervals.

AI LAB TEAM

Evaluated by an Interdisciplinary Team

5 full-time evaluation engineers and 2 visiting academic advisors, with professional backgrounds covering NLP, recommendation systems, financial risk control, and computational linguistics.

CT

Dr. Chen

Lab Director · Former BAT NLP Algorithm Expert · PhD in Computational Linguistics

LZ

Mr. Li

Evaluation System Architect · Large-scale Distributed Stress Testing · 8 Years ML Infrastructure Experience

WJ

Ms. Wang

Chinese Corpus Expert · Former Institute of Linguistics, CASS · C-Eval Co-builder

ZS

Mr. Zhang

Multimodal Direction · Vision / OCR / Chart Understanding Evaluation

OUR BOUNDARIES

What We Do and What We Don't Do

The clearer the boundaries, the more trustworthy the evaluation conclusions.

WHAT WE DO

What We Stand For

  • 01Independently purchase all model APIs for evaluation
  • 02Conduct real tests and publish raw data
  • 03Apply equal rigor and fairness to every model
  • 04Publish failure cases and negative evaluations
  • 05Quarterly review of evaluation biases
WHAT WE NEVER DO

What We Never Do

  • 01Do not accept vendor-sponsored evaluations or PR articles
  • 02Do not sign exclusive or non-compete agreements
  • 03Do not modify scores to please vendors
  • 04Do not hide any model's flaws
  • 05Do not participate in model training data collaborations

Evaluation is an attitude. If scores change due to business relationships, they are not scores—they are advertisements. What we produce are scores.

— Dianming AI Lab

CHOOSE WITH EVIDENCE

Choose the Right Model with Real Evaluation
Spend Every Budget Cent on Results

Based on unified standards of continuous real testing, providing more honest model selection criteria for different business scenarios.

Browse Model Marketplace