AI Evaluation
An AI evaluation, often shortened to an eval, is a structured test used to measure a model or system's capabilities, reliability or behavior.
Evals can measure anything from mathematical reasoning and coding to autonomy, scientific capability, safety or susceptibility to misuse.
As models become more capable, designing evaluations that continue to reveal meaningful differences becomes increasingly difficult.
← Back to Index