PRÆVIDEO Start a conversation

INDUSTRY BENCHMARKS

Test the work
you need done

A reusable benchmark for your industry helps you choose an agent, find its limits, and assess each upgrade against the same work.

Discuss your benchmark
A STANDARD FOR USEFUL WORKPVG / FIELD STUDY
Start with a job that matters to the people using your software.

Your customers
set the standard

A polished answer is not necessarily a completed job. A schedule must be feasible. A reconciliation must balance. A recommendation must respect the operator’s authority.

We define test cases with people who know the work. Your team receives the task set, scoring criteria, a record of baseline runs, and an explanation of failures. Use them to compare candidate agents, prioritize training, and check changes before exposing customers to them.

ILLUSTRATIVE TASKS

Three jobs worth testing

Field operations
Recover a disrupted schedule

A technician becomes unavailable. Given open jobs, qualifications, travel times, and service windows, propose a revised plan. Check for double-bookings, missed qualifications, and unapproved overtime. An impossible schedule should trigger an explicit escalation, not a fabricated assignment.

Download task brief

Manufacturing
Replan after a machine outage

Given job routings, material availability, machine capacity, and due dates, revise the affected production sequence. Require each operation’s dependencies to hold. Count infeasible allocations as failures; assess lateness and cost only among feasible plans.

Download task brief

Professional services
Explain a reconciliation difference

Given a ledger, statement, and supporting documents, match entries and explain the remaining difference. Require source references and arithmetic that balances. Unsupported adjustments, invented evidence, or posting without approval fail the task.

Download task brief

Task-design templates for your team to adapt. These are not benchmark datasets or performance results.

Completion, quality,
and operating cost

Report task success alongside constraint violations, escalation behavior, latency, and run cost. Keep results by task type, so strong routine performance cannot conceal a weak exception-handling path.

Compare against a model without task-specific post-training and, where practical, the existing workflow. Record human review separately. A benchmark informs deployment judgment; it does not replace it.

TECHNICAL PROTOCOL

What the benchmark specifies

The task contract

Input and output. Package the initial state, instructions, reference material, and required output schema. State whether the agent must propose an action or execute it in a test environment.

Allowed tools and actions. List tool interfaces, access boundaries, action budgets, and approval points. Fix the environment and reset it between attempts.

Success and failure. Define executable outcome checks where possible, explicit disqualifying violations, and a rubric for expert judgment. Preserve outputs and action traces for review.

The comparison contract

Held-out cases. Keep evaluation cases separate from training, including unfamiliar combinations and difficult exceptions. Document coverage and checks for overlap.

Versioning. Identify the task set, scoring rules, environment, model, prompts, and tool configuration for every run. A changed test receives a new version; rerun the baseline before comparing across versions.

Repeatability. Record run settings, retries, failures, and the number of attempts. Show variation across repeated runs where behavior is nondeterministic. Distinguish model failures from test infrastructure failures.

Bring a job
and its hard cases

Start with a workflow description and what your experts reject. We’ll scope the cases, rights, review effort, and delivery format. No customer records are needed for the first conversation.

Scope an evaluation