AI Coding & App Builders
Debug an Application From Logs Before Editing Code
Debug a reproducible failure from preserved logs using a factual timeline, ranked hypotheses, discriminating tests and regression proof.
What you will have at the end
A preserved reproduction, fact-only timeline, ranked cause table, minimal verified fix and explicit record of residual uncertainty.
- Difficulty
- Intermediate
- Time
- 45–120 minutes depending on reproducibility.
Testing scope
What was actually exercised, and what still requires verification in your own environment.
The dry run used a small synthetic repository with a deterministic API timeout and a misleading secondary error. Codex analyzed the fixture, selected a discriminating test and reviewed the patch. Cursor was not connected and no production logs or repository were accessed.
Tools referenced
Profiles and official implementation options used by this workflow; see the testing scope for which integrations were exercised.
Steps
Step 1: Preserve logs, environment and reproduction
Save the complete error output, exact command or request, version identifiers and relevant configuration with secrets redacted. Reproduce in the smallest safe environment and record whether the failure is deterministic. Keep the original log unchanged; annotate in a separate file. If reproduction is unsafe or unavailable, say so and lower confidence. A screenshot or final exception alone is not enough when earlier events may explain the failure.
Step 2: Build a factual failure timeline
Use the timeline prompt to order observable events by timestamp or causal sequence. Include successful operations immediately before the failure, primary error, retries and secondary errors. Label each row fact, calculated observation or unknown. Do not insert a root cause. Verify every fact against the preserved log, code or configuration. This discipline prevents a loud cleanup error from being mistaken for the initiating fault.
Step 3: Rank hypotheses and discriminating tests
Generate a short hypothesis list tied to timeline evidence and contradicting facts. For each hypothesis, define the smallest test whose outcomes would materially raise or lower its likelihood. Prefer inspection or a focused test over broad edits. Ask Cursor to reference files and lines, but open those locations yourself. Avoid accepting a cause simply because the model can describe it convincingly.
Tool: Cursor
Step 4: Run one small test before editing
Choose the safest, highest-information test and predict the expected result for each leading hypothesis before running it. Capture the command and output. Update the hypothesis table from the observation. If the result is ambiguous, design another test instead of editing several suspected areas. The gate passes when one cause has direct supporting evidence and plausible alternatives are either reduced or explicitly remain open.
Step 5: Implement and review the minimal fix
Change only the code necessary to address the supported cause. Add or update a test that fails for the original behavior, confirm the failure when practical, apply the fix and confirm the test passes. Run the proposed-fix review prompt on the diff, then manually inspect side effects, error handling and data boundaries. Do not combine cleanup refactoring with the diagnostic patch unless separately justified.
Step 6: Run regression checks and document uncertainty
Execute focused tests, relevant broader tests and the original reproduction. Compare logs with the baseline and confirm the misleading secondary error is understood, not merely hidden. Record untested paths, environment differences and monitoring needs. A passing test proves only its asserted behavior. Close the incident with commands, results, diff and rollback note so another engineer can verify the conclusion.
Open the tools
Official sites for implementation. Their presence here does not mean a live account or integration was tested.
Prompts used
Copy them from the linked pages.
- Build a factual failure timeline · tested on OpenAI GPT-5 (Codex) — editorial dry run
- Rank causes and discriminating tests · tested on OpenAI GPT-5 (Codex) — editorial dry run
- Review a proposed fix · tested on OpenAI GPT-5 (Codex) — editorial dry run
Editorial validation record
Illustrative scenario reviewed on Oct 3, 2026. This is not proof that the named third-party integrations were run.
- OutputIllustrative scenario: primary error isolated
The synthetic log contained a timeout followed by a cleanup exception. A focused latency test supported the timeout cause; changing the cleanup handler alone was rejected. The minimal fixture patch passed the original reproduction and focused regression check.
- DatasetTesting limitation
Only the synthetic fixture was analyzed in Codex on 2026-10-03. Cursor, a live repository, production services and production logs were not used.
Last verified
Quick answers
- How long does it take?
- 45–120 minutes depending on reproducibility.
Sources
- Cursor documentationaccessed
- Cursor security informationaccessed
More AI Coding & App Builders workflows
Prepare an AI-Built App for Human Handoff
A reproducible repository handoff with architecture and environment inventory, build/test evidence, security findings, known limitations and prioritized backlog.
Advanced · 2–4 hours for a small prototype. · Editorially reviewed Oct 3, 2026
Turn Requirements Into a Unit-Test Matrix and Tests
A requirement-to-test matrix, generated and reviewed tests, red-green execution evidence, negative cases and an explicit uncovered-requirements list.
Intermediate · 60–120 minutes for a small module. · Editorially reviewed Oct 3, 2026
Run a Risk-Based AI Pull Request Review
A prioritized review record whose findings link to requirements, changed code, evidence, verification results and a final human disposition.
Intermediate · 30–90 minutes depending on diff size and risk. · Editorially reviewed Oct 3, 2026
New workflows by email
New editorial workflows and changes to the tools they reference. Sponsored items are labelled.