Our review method: make the process and result inspectable
A transparent recording method, with one case analysis and no standardized scores or rankings.
Understand which conclusions require direct evidence.
On this page
01Current status
No standardized results, scores, or model rankings have been published. This page describes our planned recording method.
Tutorials help readers reproduce tasks; reviews record conditions and actual performance. One smooth demonstration is not a general result.
See the recorded webpage walkthrough and case review. Both preserve limitations and are excluded from model rankings.
02Record six things
- Environment: date, operating system, app version, model, and reasoning settings.
- Input: full request, starting materials, and agreed checks.
- Process: start/end times, retries, manual edits, and permission interactions.
- Artifacts: source, inspectable output, and key screenshots.
- Verification: passed, failed, and unperformed checks.
- Cost: actual visible usage and charges only; mark unavailable data explicitly.
03Classify outcomes
- Complete
- All agreed checks pass and the artifact is inspectable.
- Partially complete
- A usable artifact exists, but specific requirements are unmet.
- Incomplete
- The main goal was not achieved; record the stopping point.
- Unverified
- A necessary check was not performed, so do not label it a success or failure.
Future comparisons should use matched tasks, materials, and constraints, while disclosing differences. A single run supplies evidence only about that run.
About this article
This page describes the handbook’s review method. The existing single-case review is not a standardized benchmark or model ranking. Updated: 2026-10-02.