Skip to main content
AI & Full-Stack·16 min read

Production Evals for Coding Agents in TypeScript and Next.js (2026)

Code generation is cheap. Knowing the agent is right is not. This guide shows how to score TypeScript and Next.js coding agents with fixture tasks, deterministic checks, LLM judges used only where needed, merge gates, and a CI harness you can run without a human watching every diff.

By Mussawar Hayat

Agents write the first draft. Evals decide if it ships.

The conversation on X this week is no longer “does the agent write code?” It is “do you trust it to merge?” Coding agents plan, edit, test, and open pull requests. The bottleneck moved from keystrokes to judgment. If you cannot score an agent the same way you score a CI job, you do not have a software factory. You have a faster way to ship unreviewed mistakes.

This guide is a production eval harness for TypeScript and Next.js teams. You will define tasks with known good outcomes, run the agent against a frozen repo snapshot, score with deterministic checks first, use a model judge only for the leftover, and fail the merge when the score drops below a gate you chose in advance.

What you will implement

  • A task schema: prompt, repo fixture, allowed tools, expected files, tests
  • Deterministic scorers: typecheck, unit tests, lint, forbidden paths, secret scan
  • An optional LLM judge with a rubric, never as the only score
  • A Next.js Route Handler that records runs and refuses to “pass” incomplete evidence
  • A CI job that blocks merge when the rolling eval set regresses

1. The problem

A coding agent can look finished and still be wrong. It may add a passing test that asserts the bug. It may typecheck while changing the public API. It may “fix” a flake by deleting the assertion. A human reviewer can catch that once. They cannot catch it on every background agent that opens a PR at 2 a.m.

Chat-style vibes checks do not survive that load. You need a fixed set of tasks, a frozen environment, and scores that do not depend on the model being in a good mood.

2. What an eval is (and is not)

An eval is a repeatable experiment: same prompt, same starting tree, same tool budget, scored by the same functions. It is not a demo, a leaderboard screenshot, or a single lucky green build.

Split scores into two families.

  1. Deterministic. tsc --noEmit, vitest or playwright, ESLint, ripgrep for forbidden strings, file-existence checks, git diff path allowlists. These should carry most of the weight.
  2. Judged. A second model grades design notes, comment quality, or whether the PR description matches the diff. Use this only when a test cannot express the property. Never let the same model family grade its own work without a rubric and a human-audited sample.

Official guidance on tool use and structured outputs lives in the current model provider docs. For Next.js itself, treat the App Router, Route Handlers, and generateStaticParams the way the Next.js documentation describes them — do not invent APIs in the eval fixtures.

3. Task format

Store tasks as JSON or TypeScript objects in the repo, not in a chat log. Each task is versioned. Changing the prompt or the fixture bumps the task id so historical scores stay comparable.

export type AgentTask = {
  id: string
  version: number
  prompt: string
  fixtureDir: string
  allowedPaths: string[]
  forbiddenPaths: string[]
  maxSteps: number
  checks: {
    typecheck: boolean
    testCommand: string
    mustExist: string[]
    mustNotContain: { file: string; pattern: string }[]
  }
  judgeRubric?: string
}

Example task: “Add a rate-limited POST /api/quotes Route Handler that validates a Zod body and returns 429 with Retry-After.” The fixture is a small Next.js app with no handler. The tests already exist and fail. The agent may only write under app/api/quotes and lib/rate-limit.ts.

4. Runner

The runner is a Node script, not a chat window.

  1. Copy the fixture into a throwaway worktree.
  2. Start the agent host with a tool budget and a wall-clock timeout.
  3. Allow only the tools you would allow in production: read, edit inside allowedPaths, run the documented test command. No unrestricted shell. No network except the model API.
  4. Collect artifacts: git diff, test stdout, step log, token counts if the host exposes them.
  5. Score. Persist a row. Delete the worktree.

Do not reuse a dirty workspace. Yesterday’s half-applied patch is a hidden variable.

5. Deterministic scorers

Write scorers as pure functions over artifacts. Each returns { pass: boolean; detail: string; weight: number }.

  • Typecheck. Fail the task if tsc is non-zero, even if tests passed by using any.
  • Tests. Run the exact command stored on the task. Do not let the agent rewrite the test file unless the task says so. If it may rewrite tests, add a second frozen oracle test that the agent cannot touch.
  • Path policy. Reject diffs that touch forbiddenPaths or leave the allowlist.
  • Secret scan. Fail on patterns that look like keys, .env dumps, or copied production URLs.
  • Invariant strings. For API tasks, assert response shape with a recorded HTTP fixture against the built route, not by reading the source and hoping.

Weighted sum is enough. You do not need a research paper. You need a number that moves in the right direction when the agent actually improves.

6. When to use an LLM judge

Use a judge for properties tests cannot see: “the error message is usable,” “the PR description lists the breaking change,” “the component is accessible enough to review.” Give the judge the rubric, the diff, and the test output. Do not give it the agent’s self-grade.

Pin the judge model and temperature. Store the raw judge JSON. Sample 10 percent of judged tasks for a human to relabel every sprint. If human agreement falls, the rubric is the bug, not the agent.

7. Next.js recording API

Expose POST /api/evals/runs behind an internal token. The body is the task id, git sha of the harness, agent name, scores, and artifact hashes. The handler validates with Zod, writes with Prisma, and never accepts a client-supplied passed: true without recomputing pass from the score vector on the server.

Read path: GET /api/evals/tasks/:id/history returns the last N runs so a dashboard or a CI annotation can show regressions. Keep tenant or repo id on the session, not in the body, same rule as any other write gateway.

8. CI merge gate

Pick a small golden set — 15 to 40 tasks that represent the work you actually assign agents: Route Handlers, Prisma migrations you already reviewed, Tailwind refactors, test-only changes. Run that set on every candidate agent config and on a nightly job against main.

Gate rule example: mean score must stay within 3 points of last week’s baseline, and no golden task may flip from pass to fail without a harness version bump. When a task is too easy, replace it. An eval set that always scores 100 is decoration.

Do not block product PRs on a flaky judge. Block them on deterministic failures. Put judge drift on a dashboard.

9. Security

  • Fixtures must not contain production secrets or real customer data.
  • The agent sandbox must not reach your staging database.
  • Judge prompts must not include other tenants’ diffs.
  • The recording API is an internal write surface: auth, rate limit, allowlisted fields.
  • Treat eval prompts as code. Review them. Prompt injection in a fixture can teach the agent a bad habit you then promote.

10. Cost and performance

Full-repo agents are expensive. Keep fixtures small. Cap steps. Cache model embeddings only if you are doing retrieval, not if you are scoring patches. Run deterministic checks locally in CI; run judged tasks on a schedule. Record token use per task so a “smarter” agent that costs 8x for +1 point is a visible tradeoff, not a surprise invoice.

11. Real use cases

  • Choosing between two coding hosts on your codebase instead of on a public leaderboard.
  • Catching a prompt or skill change that made the agent start rewriting lockfiles.
  • Teaching a new hire which tasks are safe to hand an agent unsupervised.
  • Proving to a client that an automated PR bot is measured, not hoped.

12. Common mistakes

  • Scoring only “did tests pass” after the agent was allowed to edit the tests.
  • Using the same model as judge and implementer with no rubric.
  • Changing the fixture and keeping the old task id.
  • Running evals on a dirty worktree.
  • Publishing pass rates without the harness version.
  • Treating one viral benchmark as a substitute for tasks that look like your app.

13. FAQ

Can I skip unit tests and only use an LLM judge?

No. Judges are slow, expensive, and drift. Tests and typecheck should decide most of the score.

How many tasks do I need?

Start with a golden set you can run in one CI hour. Quality of fixtures beats a 500-task set you never re-read.

Should the agent see the hidden tests?

No. Public tests can guide it. Oracle tests stay out of the worktree.

Do I eval the host or the model?

Both, but change one variable at a time. Record host, model id, skill pack, and harness version on every run.

When is a human still required?

Architecture, threat models, irreversible migrations, and anything that can move money or personal data. Evals shrink the review surface. They do not retire ownership.

14. Summary

Coding agents made generation cheap. Production still cares about typecheck, tests, blast radius, and who debugs the failure. An eval harness turns that concern into a number you can gate on: frozen fixtures, deterministic scorers first, judged properties second, recorded runs, and a CI rule that fails when the agent quietly gets worse.

Key takeaway

If you cannot rerun last week’s tasks on this week’s agent and get a comparable score, you are not evaluating. You are watching a demo.


Need an eval gate on a Next.js / TypeScript agent stack?

I build agent write paths, MCP servers, and CI scoring for teams that have to defend automated PRs. Get in touch or see full-stack and AI development services.

Related reading: Postgres write guardrails for agents and OpenAI Agents SDK multi-agent workflows.