Pilot evaluation
Measure whether the pilot improves human review
Use a small scorecard to evaluate review clarity, evidence quality, friction, and unresolved outcomes without overstating the result.
Gated · · Private-preview evaluation guide
Choose measures before seeing results
A pilot should inform a decision your team will actually make. Choose a few measures before the first request: whether a reviewer identifies the exact inputs, how much clarification is needed, whether the final provider state is reconciled, and how much effort the workflow adds. Record the current process as a baseline when a comparable example is available.
Gated’s private preview evaluates one uniquely named GitHub branch at an exact existing commit in a selected non-critical repository. Keep the scorecard confined to that workflow. Evidence about reviewer usability does not establish protection for other tools, repositories, or operations.
Count outcomes with clear definitions
Define a completed live operation as one whose execution evidence and independent provider readback have been reconciled with the reviewed repository, branch, and SHA. Track approvals separately. An approved request awaiting explicit execution is not a completed write, and a lost response is not automatically a failure or success.
Track safe-mode rehearsals in a separate group. They can reveal review friction and missing context, but cannot count as live GitHub writes from the runtime. Retried calls for the same request should stay associated with that request rather than inflating the number of independently reviewed operations.
- Review clarity: exact inputs understood, clarification needed, or unresolved.
- Evidence: provider state matched, mismatched, absent, or unknown.
- Effort: time to review and time spent reconciling an outcome.
- Boundary: alternate credentials known, resolved, or still unexplained.
Preserve the cases that are uncomfortable
Include denied requests, abandoned reviews, mismatched evidence, and uncertain outcomes in the record. Describe why a request did not proceed, using the observed reason rather than guessing. A small number of successful runs can be useful, but should not erase cases that exposed confusing wording or an unclear owner.
Use the same definitions throughout the evaluation period and record any change. Keep sensitive evidence in an appropriate team location; share only redacted summaries externally. Avoid percentages without the underlying counts, especially when the pilot contains only a few requests.
Make the next decision specific
Review the scorecard with the person who approves requests and the repository owner. Identify whether the workflow made decisions clearer, whether the execution outcome was understandable, and where added effort was acceptable. Assign concrete follow-up work to unanswered questions before changing the scope.
Conclude with a decision to continue the same narrow pilot, address a named issue, or stop evaluation. These observations do not prove security effectiveness, production readiness, or universal tool interception. Direct credentials can bypass the controlled path, and merge, deletion, and deployment governance remain outside the supported pilot.
Evaluate one controlled GitHub workflow
Read the verified GitHub scope and pilot entry steps. Join the update list if you want to follow the preview. Signup does not grant immediate access or commit you to a purchase.