Orcha
Tutorial and Examples

Evaluate an agent

Add OrchaJS evaluations that score completed agent responses with a model judge.

evaluate-an-agent~/evaluate-an-agent/orcha/writingAgent/evaluations/responseQuality/index.json
01{02  "name": "response_quality",03  "description": "Clarity and accuracy of explanations.",04  "instructions": "Judge the complete answer using the session evidence.",05  "provider": "openai",06  "model": "gpt-5-mini",07  "metrics": [08    {09      "name": "clarity",10      "description": "The answer explains the concept directly.",11      "threshold": 0.812    },13    {14      "name": "calibration",15      "description": "The answer avoids unsupported claims.",16      "threshold": 0.817    }18  ]19}

Measure quality, not exact wording

This example is a complete writing agent with an asynchronous model judge. The agent answers with Anthropic, while the response_quality evaluation uses OpenAI to score clarity and calibration independently.

Evaluations are for qualities that cannot be captured reliably by substring or exact-value assertions. Every metric has a concrete criterion and an inclusive threshold from 0 to 1.

Judging starts after a run completes. execution.result becomes available without waiting, while execution.evaluations resolves when every enabled judge has finished and its result has been stored.

Select any file in the editor to inspect the complete standalone project.

Understand each file

orcha/writingAgent/evaluations/responseQuality/index.json

Configures an asynchronous model judge and its quality thresholds.

  • The judge can use a different provider and model from the agent.
  • Each metric receives a score from 0 to 1.
  • A metric passes when its score meets or exceeds threshold.
orcha/writingAgent/index.json

Configures the agent whose completed responses are evaluated.

orcha/writingAgent/instructions.md

Defines the behavior the evaluation metrics are intended to measure.

orcha/index.ts

Configures both the agent provider and the independent judge provider.

app.ts

Awaits the agent result and the separate asynchronous evaluation promise.

  • execution.result does not wait for judges.
  • Await execution.evaluations when scores must finish before the process exits.

Next: deploy an agent.