PlatformBenchmarksModelsComparisonAboutOpen app
A Puma experiment
getEvals

Useful evals for real-world tasks.

A Puma experiment. Measure AI on work that matters to you — privately, on your device — not someone else's leaderboard.

Private by design

Public benches show how models behave in the wild. Private evals tell you whether your prompts, tools, and agents hold up on the tasks you actually run. Your gold answers stay on your device.

Cases you keep

Capture the tickets, documents, and workflows your system must handle. Nothing is uploaded. You are not training the industry on your test set.

Graders you control

Score with exact match, contains, regex, JSON field checks, and keyword lists. Start simple, then tighten the rubric as the work gets clearer.

Runs that stay local

Paste model outputs, label the system under test, and keep a history of pass rates as prompts and models change — all in this browser.

Two paths. One purpose.

getEvals sits next to Puma Browser and the PumaAI lab. Same purpose: AI that stays in your hands.

Start with the demo suite Browse public benchmarks