Factories > Measure & improve
Configuring and reviewing Self-improvement
# Configuring and reviewing Self-improvement {/* VISUAL: The Self-improvement pull request list, or a Benchmarks suite run -- this section is text-only today. */} Turn on **Self-improvement** for each Scorer whose failures you want investigated automatically. A failure is a score below the Scorer's pass threshold. By default, the scheduled check starts a Self-improvement run after an agent has 25 distinct unreviewed failures, or when its oldest unreviewed failure is seven days old. It groups the failures for each agent into a follow-up run that proposes a fix. For a GitHub-backed factory, set `selfImprovement.failedRunThreshold` in `factory.yaml` to change the scheduled threshold to any value from 1 through 50. The `reviewerType` values are `admins` (the default, team admins and owners), `team` (any team member), `none`, and `custom` (members listed in `reviewerEmails`). Warp randomly requests one eligible reviewer unless you use `none`. See [`selfImprovement` in the factory definition reference](/factories/factory-as-code/#selfimprovement). To run the check without waiting for the scheduled threshold, click **Run now** on the factory dashboard's **Self-improvement** page. An ad hoc run can include an agent with one unreviewed failure. You cannot choose which agents or failures it processes. A follow-up run can propose changes to application code. It can also improve the factory itself: when you manage your factory as [definitions as code](/factories/factory-as-code/), its prompts, skills, and configuration are version-controlled files, so a follow-up run can open a pull request against the factory definition the same way it would against application code. Nothing is adopted without your review. The factory dashboard's **Self-improvement** page lists the pull requests these follow-up runs open. Each pull request includes a **Regressions addressed** section that links the failing runs and Scorer results behind it, so you can trace the change back to its evidence. ## Related pages * [Measure and improve a factory](/factories/measure-and-improve/) - Read dashboard metrics and run a practical improvement loop. * [Configuring Scorers](/factories/measure-and-improve/scorers/) - Define the criteria whose failures Self-improvement investigates. * [Benchmarking factory agent configurations](/factories/benchmarks/) - Compare model and runner configurations on the same tasks.Tell me about this feature: https://docs.warp.dev/factories/measure-and-improve/self-improvement/Turn on Self-improvement for a Scorer to turn its repeated failures into follow-up pull requests you review.
Turn on Self-improvement for each Scorer whose failures you want investigated automatically. A failure is a score below the Scorer’s pass threshold. By default, the scheduled check starts a Self-improvement run after an agent has 25 distinct unreviewed failures, or when its oldest unreviewed failure is seven days old. It groups the failures for each agent into a follow-up run that proposes a fix.
For a GitHub-backed factory, set selfImprovement.failedRunThreshold in factory.yaml to change the scheduled threshold to any value from 1 through 50. The reviewerType values are admins (the default, team admins and owners), team (any team member), none, and custom (members listed in reviewerEmails). Warp randomly requests one eligible reviewer unless you use none. See selfImprovement in the factory definition reference.
To run the check without waiting for the scheduled threshold, click Run now on the factory dashboard’s Self-improvement page. An ad hoc run can include an agent with one unreviewed failure. You cannot choose which agents or failures it processes.
A follow-up run can propose changes to application code. It can also improve the factory itself: when you manage your factory as definitions as code, its prompts, skills, and configuration are version-controlled files, so a follow-up run can open a pull request against the factory definition the same way it would against application code. Nothing is adopted without your review.
The factory dashboard’s Self-improvement page lists the pull requests these follow-up runs open. Each pull request includes a Regressions addressed section that links the failing runs and Scorer results behind it, so you can trace the change back to its evidence.
Related pages
Section titled “Related pages”- Measure and improve a factory - Read dashboard metrics and run a practical improvement loop.
- Configuring Scorers - Define the criteria whose failures Self-improvement investigates.
- Benchmarking factory agent configurations - Compare model and runner configurations on the same tasks.