Skip to content

SEP-2000: Expose a run's failure reason on the task history payload - #1473

Merged
yyyyyyyan merged 7 commits into
mainfrom
SEP-2000
Sep 8, 2026
Merged

yyyyyyyan merged 7 commits into
mainfrom
SEP-2000

Conversation

@yyyyyyyan

@yyyyyyyan yyyyyyyan commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A failed run now says why on the task-history payload, so an operator does not have to open its log to learn anything.

Adds a nullable failure_reason column to TaskHistory via one additive migration on the tasks Alembic track, with no backfill. It is declared on TaskHistoryBase, so it reaches the table model and TaskHistoryResponse together and flows through the Tasks API list/retrieve payloads and the SEP proxy that re-serves them with no serialization change.

Two single points of enforcement:

  • TaskHistory.set_failure_reason() is the only write path. It collapses whitespace runs (including newlines) to single spaces, maps a blank value to None rather than "", and truncates through shorten_text at MAX_FAILURE_REASON_LENGTH (500). The column is declared as plain str/TEXT rather than VARCHAR(500) deliberately: PostgreSQL enforces a declared length and SQLite does not, so a bound declared in the schema would fail in production in a way the test suite structurally cannot observe. The bound is enforced at write time and pinned by a test instead.
  • TaskHistoryStatusEnum.operator_summary() is the only prose source. The four operator-facing fragments move off alert_for_status onto the enum, and both the alert summary and the stored reason now read them from there, so the two cannot drift. The four strings are byte-identical to the literals they replace; alert_for_status's dedup keys, severities, alert classes, SUCCESS resolve arm and final payload are untouched, and its existing tests pass unmodified.

Reasons are composed only from structural facts SEP already owns:

Failure site Reason
_persist_failed_dispatch (unhealthy target, unresolvable payload) the string it already receives and writes to stderr — now persisted too
Nomad, failure with an allocation the failing producing step and the exit code of its last termination
Nomad, failure whose allocation names no failed producing step none — the row reports null, meaning "unknown"
Nomad, failure with no allocation canned prose naming the missing allocation
Nomad lost / stale / unlaunchable the status's own prose
Celery executor the exception type, plus the dotted callable path when it is a string inside the allowed namespace

The exit code is read off the step's last Terminated event, so a step Nomad restarted reports the termination the allocation ended on rather than its first attempt's. The failing step itself is chosen by execution order, through the same _alloc_step_sort_key the module already uses, not by the order Nomad serialized the task states in. Nomad's JSON encoder emits map keys sorted, so picking the first key would report Step 'clean-up' failed for a run whose payload failed and whose cleanup then failed after it — blaming SEP's own cleanup for the payload's failure.

The enum's fragments are verb phrases sized for alert_for_status's mid-sentence slot, so the stored reason renders them as standalone sentences (Execution tracking lost.) rather than giving them a subject — The run execution tracking lost. is not a sentence.

The Celery path takes the exception type and callable path rather than str(exc). That path resolves an arbitrary dotted callable from task.data["callable"] and runs it in-process, so a rendered exception can carry a DSN with credentials or quoted row data, and this field has no masking. Task.data is a raw JSON column with no validator behind it, so the callable path itself is echoed only when it is a str under CELERY_CALLABLE_ALLOWED_PREFIX; any other stored value composes the generic Task execution raised <Type>. The traceback still goes to the stderr log unchanged. For the same reason the Nomad composer reuses only the step name and exit code, never the Nomad event's own description text, which can carry a driver-supplied exit message.

_apply_terminal_status sets a reason on every arm, None included, so "a non-failed status carries null" is an invariant of the writer rather than an absence. BaseExecutor.stop_task clears the reason only on its forced-STOPPED arm — a run the sync had already resolved as failed keeps the reason explaining why, which is what that method's docstring requires.

Also in this diff: POST /history/ is repaired

Unrelated to the feature, found while covering the new normalization on that write path. The route could not have served a request on main:

  • It logged task.name on a TaskHistory, which has no name attribute — AttributeError on every call.
  • Past that, its response_model requires task and execution_request, and TaskHistoryManager.save re-defers execution_request, so serializing the save's own return value attempts lazy IO from an async context (MissingGreenlet).

It now logs task_id and re-reads the saved row with task joined and execution_request undeferred. Two tests cover it, and it ships its own fixed changelog fragment. No in-repo caller uses this route — every other /history/ reference is a GET — which is why the breakage went unnoticed.

No authorization or identity change anywhere in this diff. create_task_history keeps IsAuthenticatedDep; no guard, principal or role requirement is added, removed or altered. An authenticated caller can still supply arbitrary text in failure_reason — normalization bounds the shape, not the content — exactly as that route already accepts execution_request, status and executed_by. Tightening the create contract is a separate concern.

Historic rows read NULL, and a NULL on a failed row means "unknown", not "did not fail". The migration docstring and the changelog fragment both say so.

Scope note

The route repair above is delivered beyond what any acceptance criterion names; it contradicts none of them. Recorded here rather than left silent so the ticket and the diff do not disagree.

Tested

  • Run a Nomad-backed task whose script exits non-zero; confirm the history row reports the failing step and its exit code, and that the same run's stderr log still holds the full output.
  • Run a Nomad-backed task against a target host that is not ready; confirm the history row's reason matches the text written to the run's stderr log.
  • Run a task whose file:// payload reference is missing; confirm the history row reports that the payload could not be resolved and names the task.
  • Run a Celery-backend task whose callable raises; confirm the reason names the callable's dotted path and the exception type, and that neither the exception message nor the traceback appears in the field while both remain in the stderr log.
  • Let a task be skipped as stale, and separately be reported unlaunchable; confirm each carries its own prose and that the PagerDuty alert summary for the same run reads exactly as it did before this change.
  • Complete a task successfully, and separately stop a running one; confirm both report no reason.
  • Stop a task that has already failed; confirm the failure reason survives the stop rather than being cleared.
  • Open a task-history list containing a run recorded before this release; confirm the field is present and null rather than absent.
  • Fetch task history through the SEP proxy both with and without task_names (the merge and passthrough paths); confirm the reason appears on both.

Checklist

  • New/modified functions have type hints and rST docstrings
  • New tests added for new features or bug fixes
  • Database migrations generated if models changed (make makemigrations)
  • User-facing changes documented (README, inline help, UI text)
  • Configuration changes documented with examples

Add a nullable, operator-facing `failure_reason` to `TaskHistory`, populated
at each failure site that already determines a reason and serialized on
`TaskHistoryResponse`, so a failed run says why without the operator opening
its log.

`TaskHistory.set_failure_reason` is the single write path: it collapses a
reason to one line and bounds it at `MAX_FAILURE_REASON_LENGTH`, mapping blank
to `None`. `TaskHistoryStatusEnum.operator_summary` is the single prose source,
so the alert summary and the stored reason cannot drift.

Reasons are composed from structural facts only — step name and exit code for
Nomad, exception type and resolved callable path for Celery, canned prose for
lost / stale / unlaunchable. The traceback and the executor's own output keep
going to the stderr log and are never used as the reason.

Also repair `POST /history/`, which could not have served a request: it logged
`task.name` on a `TaskHistory` (`AttributeError`), and its response model
needed `task` and `execution_request`, which `save` re-defers. It now re-reads
the saved row with both loaded. Found while covering the new normalization on
that write path.

Existing rows read NULL and are not backfilled, so a missing reason means
"unknown", not "did not fail".
Read the failing producing step and its state in one pass over the
allocation's task states, rather than re-reading the whole mapping per
step through `_alloc_step_state`.

Drop the optional `:type` directives from `TaskHistoryBase`'s docstring —
the annotations are the source of truth — and promote the duplicated
`tasks_alembic_config` fixture into a `conftest.py` for the Tasks-track
migration tests.

Give the downgrade test a positive control so it cannot pass by asserting
absence against a table that is not there.
`_failed_step_reason` returned the first failed producing step in the
allocation's serialization order. Nomad's JSON encoder emits map keys
sorted, so a run whose payload failed and whose cleanup then also failed
was reported as `Step 'clean-up' failed` — blaming SEP's own cleanup for
the payload's failure. Walk the steps through `_alloc_step_sort_key`, the
ordering the module already uses elsewhere, so the earliest failure wins.

Render the enum's status prose as a standalone sentence rather than
prefixing it with "The run". The fragments are verb phrases written for
the mid-sentence slot in `alert_for_status`, where "execution tracking
lost" reads correctly; "The run execution tracking lost." does not.

Clear the reason on the Celery executor's success arm too, so "a
non-failed status carries no reason" holds at every terminal writer
rather than only the Nomad ones.

Add the changelog fragment owed for the `POST /history/` repair, and
correct the bound constant's comment, which claimed every composed reason
sits well under it — the payload-resolution reason embeds a filesystem
path and is unbounded when composed.
Copilot AI balanced review requested due to automatic review settings September 7, 2026 18:51
@yyyyyyyan yyyyyyyan added the qa in progress Someone is currently testing this PR - do not merge it label Sep 7, 2026
@yyyyyyyan
yyyyyyyan requested a review from a team as a code owner September 7, 2026 18:51
@yyyyyyyan yyyyyyyan self-assigned this Sep 7, 2026
@github-actions github-actions Bot added python frontend svc:tasks PR touches the tasks service (app/tasks/) labels Sep 7, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Failure-reason semantics permit contradictory records and potentially expose unvalidated callable data.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Adds normalized failure reasons to task-history payloads, backed by an additive Tasks migration and comprehensive executor/API coverage. It also repairs task-history creation serialization.

Changes:

  • Persists bounded failure reasons for Nomad, Celery, and dispatch failures.
  • Exposes the field through Tasks and SEP APIs.
  • Repairs POST /history/ and updates schemas, clients, tests, and changelogs.
File summaries
File Description
app/tasks/models.py Defines failure-reason storage and summaries.
app/tasks/routes.py Persists and returns normalized reasons.
app/tasks/celery.py Saves pre-dispatch failure reasons.
app/tasks/execution/models.py Clears reasons on forced stops.
app/tasks/execution/executors/nomad/models.py Composes Nomad failure reasons.
app/tasks/execution/executors/celery/models.py Composes Celery failure reasons.
app/tasks/migrations/versions/2026_09_07_1500-c4b8e1f7a2d9_add_taskhistory_failure_reason.py Adds the nullable column.
tests/app/tasks/test_routes.py Covers Tasks API serialization and creation.
tests/app/tasks/test_models.py Covers normalization and status summaries.
tests/app/tasks/test_celery.py Covers persistence through worker flows.
tests/app/tasks/migrations/conftest.py Provides Tasks migration fixtures.
tests/app/tasks/migrations/test_taskhistory_failure_reason.py Tests upgrade and downgrade.
tests/app/tasks/execution/test_models.py Tests stop behavior.
tests/app/tasks/execution/executors/nomad/test_models.py Tests Nomad reason composition.
tests/app/tasks/execution/executors/celery/test_models.py Tests Celery reason composition.
tests/app/sep/api/routes/test_task_history.py Tests SEP proxy propagation.
frontend/packages/api/specs/tasks.json Updates Tasks OpenAPI schema.
frontend/packages/api/specs/sep.json Updates SEP OpenAPI schema.
frontend/packages/api/src/generated/tasks.ts Regenerates Tasks types.
frontend/packages/api/src/generated/sep.ts Regenerates SEP types.
changelog.d/SEP-2000.added.md Documents failure reasons.
changelog.d/SEP-2000.fixed.md Documents the create-route repair.
Review details
  • Files reviewed: 20/22 changed files
  • Comments generated: 3
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread app/tasks/execution/executors/celery/models.py Outdated
Comment thread app/tasks/execution/executors/nomad/models.py
Comment thread app/tasks/routes.py
@yyyyyyyan

yyyyyyyan commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor Author

Follow-up — out of scope for this PR, surfaced while reviewing it.

CeleryExecutor._run_callable resolves and invokes the dotted path in task.data["callable"] without applying CELERY_CALLABLE_ALLOWED_PREFIX. That check lives in CeleryExecutor.validate_job, which is reached only from POST /transform-payload — not from task creation and not from execution. TaskWrite accepts backend, protected and data from the request body, and TaskBase.validate_data_for_backend requires only that a Celery-backend task be protected, so nothing between the write and the invocation constrains the namespace. Both POST / and POST /execute/{task_name} carry IsAuthenticatedDep with no role floor.

SEP-1656 established HOOK_MODULE_ALLOWLIST for alert_detail_builder and run_result_recorder; the executor's own callable resolution was not part of that change.

Traced statically, not executed. Nothing in this PR touches _run_callable, validate_job or the create route's contract, so it is deliberately not being fixed here — recorded so it gets its own ticket rather than riding along.

@yyyyyyyan yyyyyyyan added qa passed Tests for this PR are completed and successful. and removed qa in progress Someone is currently testing this PR - do not merge it labels Sep 8, 2026
@yyyyyyyan

yyyyyyyan commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Automated QA — PASS

Verified the Tested items against a fresh instance on the PR head, exercising the Tasks execution framework end to end with a real Nomad node and Celery workers.

  • A Nomad task whose script exits non-zero reports the failing step and its exit code, and the full stderr stays in the run's log.

    Nomad script exit non-zero

  • A Nomad task dispatched to a target that isn't ready records a reason that is byte-for-byte identical to what's written to the run's stderr log.

    Unhealthy target reason matches stderr

  • A task whose file:// payload reference can't be resolved fails immediately, naming the task and the unresolved reference — no Nomad dispatch is attempted.

    Missing file:// payload

  • A Celery-backend task whose callable raises records the callable's dotted path and the exception type, while the exception message and traceback stay out of that field and only appear in the stderr log.

    Celery callable raises, message masked

  • An unlaunchable run (executor node can't resolve the requested command) reports its own operator-facing prose.

    Unlaunchable run reports its own prose

  • A successful run reports no reason

    Successful run has no reason

  • and a task stopped while it's still running reports no reason either.

    Task stopped while running has no reason

  • Stopping a task that has already failed on the executor side keeps the failure reason instead of clearing it.

    Stop of an already-failed task preserves the reason

  • A run recorded before this release shows the field as present and null, not missing, on the detail endpoint

    Historic row detail: field present, null

  • and on the list endpoint too.

    Historic row list: field present, null

  • The SEP gateway's task-history endpoint carries the reason on the plain list (no task_names)

    SEP proxy passthrough carries the reason

  • and on the task_names merge path.

    SEP proxy merge path carries the reason

  • A Nomad run that ends up with no allocation and no pending evaluation carries the dedicated canned reason rather than a generic failure.

  • POST /history/ — previously an unconditional AttributeError on every call — now returns the created row with its task and execution request already joined, and a caller-supplied reason is collapsed to one line and bounded the same way SEP's own reasons are.

    POST /history/ repaired, reason normalized

Observations — pre-existing, out of scope

  • The periodic-dispatch path (prepare_periodic_task_history in app/tasks/celery.py) resolves a Nomad task's target from the task's static job-template constraint rather than the caller-supplied execution_data.meta.target, so a per-invocation target override is silently ignored on that path. Pre-existing behavior, unrelated to this change.

…llers

The helper's explicit TaskHistory return type sharpened inference on
these call sites, surfacing that BaseSQLModel.id is int | None while
both the helper and maybe_record_run take int. Every id here comes off
a row that is already persisted, so the cast records that invariant
rather than widening either contract.
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown

Coverage report

Click to see where and how coverage changed

FileStatementsMissingCoverageCoverage
(new stmts)
Lines missing
  app/tasks
  celery.py
  models.py 208
  routes.py 690, 695, 722
  app/tasks/execution
  models.py
  app/tasks/execution/executors/celery
  models.py
  app/tasks/execution/executors/nomad
  models.py 476
  app/tasks/migrations/versions
  2026_09_07_1500-c4b8e1f7a2d9_add_taskhistory_failure_reason.py
Project Total  

This report was generated by python-coverage-comment-action

@yyyyyyyan
yyyyyyyan merged commit 264ce6c into main Sep 8, 2026
41 of 45 checks passed
@yyyyyyyan
yyyyyyyan deleted the SEP-2000 branch September 8, 2026 03:23
@yyyyyyyan

Copy link
Copy Markdown
Contributor Author

Tracked 2 follow-ups from this PR:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

frontend python qa passed Tests for this PR are completed and successful. svc:tasks PR touches the tasks service (app/tasks/)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants