Skip to main content
Opik Documentation

Search documentation

Type to search this documentation.

On this pageOverview

Optimizer benchmarks

We regularly evaluate every optimizer against shared datasets so you can make informed trade-offs. This page summarizes the latest results and explains how to reproduce them with the public benchmark scripts.

Each run uses Opik datasets backed by open-source corpuses commonly used in academia:

Dataset Description Primary metrics
Arc (ai2_arc) Multiple-choice science questions. LevenshteinRatio, accuracy.
GSM8K (gsm8k) Grade-school math word problems. Exact match, custom math verifier.
MedHallu (medhallu) Medical Q&A with hallucination checks. Hallucination, AnswerRelevance.
RagBench (ragbench) Retrieval-oriented questions. AnswerRelevance, contextual grounding.

Results shown below use openai/gpt-5-nano for evaluation on non multi-hop based runs. Scores will change if you select different models, metrics, agent configurations or prompt seeds.

Rank Algorithm/Optimizer Avg. Score Arc GSM8K RagBench
1 HRPO 67.83% 92.70% 28.00% 82.80%
2 Few-Shot Bayesian 59.17% 28.09% 59.26% 90.15%
3 Evolutionary 52.51% 40.00% 25.53% 92.00%
4 MetaPrompt 38.75% 25.00% 26.93% 64.31%
5 GEPA 32.27% 6.55% 26.08% 64.17%
6 Baseline (no optimization) 11.85% 1.69% 24.06% 9.81%
  1. Install dependencies (ideally in a virtualenv):
    Bash
    pip install -r sdks/opik_optimizer/benchmarks/requirements.txt
  2. Configure provider keys (e.g., OPENAI_API_KEY).
  3. Execute the runner:
    Bash
    python sdks/opik_optimizer/benchmarks/run_benchmark.py \
      --model openai/gpt-5-nano \
      --output results.json
  4. Inspect the JSON or load it into a notebook to compare against the published table.

The script spins up datasets defined in sdks/opik_optimizer/benchmarks/config.py, runs each optimizer with consistent trial budgets, and logs runs to Opik so you can review traces. Note that production use should include separate validation datasets to prevent overfitting—see Define datasets for guidance.

  • Learn how each optimizer works in the Algorithms overview.
  • Customize the benchmark configs (datasets, metrics, budgets) to mirror your production workload.
  • Share results or contribute improvements via GitHub.
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu