gpt.college

Methodology

Rankings are built from structured model, benchmark, result, and source records.

Data Source Types

Source records are labeled as official, third-party, paper, leaderboard, community, or unknown.

Production claims should prefer exact model-version sources and should separate official self-reported results from third-party evaluations.

Ranking Rules

Single-benchmark rankings sort scores according to benchmark direction. Category pages aggregate included normalized scores and ignore missing data rather than treating it as zero.

Evidence Records

Benchmark result rows identify exact model versions, source type, verification status, update date, and notes about harness or methodology differences.

Limitations

Benchmark results are useful signals, but they do not fully represent every real-world use case. Prompting, tools, sampling, model version, and evaluation harnesses can all change results.