Catch AI Regressions before They Ship with AI Evals in CI/CD
Harness, Wednesday, September 2nd, 2026
Harness AI Evals tests agent quality inside CI/CD using golden datasets and quality gates before production.
Harness introduced AI Evals, which brings AI agent quality testing into the CI/CD pipeline rather than leaving it to manual review.
The approach uses golden datasets as reference behavior and quality gates that block a release when an agent regresses against them. This targets behavioral regressions, which are harder to catch than functional bugs because an agent can still return plausible output while performing worse.
The goal is to make agent quality a pipeline concern with the same rigor as unit and integration tests.