fix(agent,tools): stop the loop reporting work it never did - #3972
kovtcharov-amd wants to merge 6 commits into
Conversation
…it use the tools it has An agent run could end with "Task completed" having written nothing, and the caller had no way to tell that from a real answer. Six defects behind that, each found by running real tasks to completion and reading what the agent actually did: - the guard that re-prompts for a missing output file called a symbol nobody imported, and a blanket except swallowed the NameError -- it had never run - a repeated tool call ended the turn claiming success; it now asks for a different approach once, then says the work is unfinished - search_file could not find a file named exactly (ci.log, pyproject.toml): an extension allowlist hid it, and the agent read the empty result as an empty workspace - content search skipped .toml, .cfg, .ts and every extensionless file, so 'bump the version wherever it is declared' never saw pyproject.toml - an unpaired </think> left a reasoning model's deliberation in the answer - the shell refused linters it already ran via 'python -m', and its hint claimed only read-only commands were allowed long after that stopped being true -- so the agent stopped trying to verify its own work Also adds the task-execution harness these were found with: real agent loop, isolated sandbox per attempt, correctness decided by a verifier's exit code and never blended with judged quality. Measured across three-run batches on the 27-task suite, shell refusals fell from 55-76% of calls to 6-10%, and accomplished rose from 19/27 to 26/27. Issues: #3941 #3942 #3943 #3945 #3946 #3957 #3967 #3971
Skill audit
✅ All audited skills cleared the tier they claim. Per-finding detail is withheld here on purpose. Read it in the Security > Code scanning tab, or download the |
Request changesThis PR bundles a large new eval-dataset subsystem ( The diff undoes recent work already on main — including a security fix. As given, it reverts the change that stopped raw OAuth provider response bodies reaching logs and user-facing errors (those responses answer requests carrying an authorization/device code), the email "a mailbox failed during the scan" caveat, the daemon 🔒 SECURITY CONCERN: the reverted OAuth error-body handling reintroduces a fixed credential-adjacent disclosure path, and separately the shell allowlist is widened in ways that need a maintainer decision. @kovtcharov-amd — please weigh in on both before this merges. Details are in the collapsed section; no exploit specifics here. The model-loading change has a bug that fires on every ordinary run. Context pinning is now off by default, but the "is it already loaded at a big enough window?" check still compares against the (now absent) expected size. The comparison throws, gets swallowed by a catch-all, and the result is that the optimisation added to avoid redundant model loads never runs — every request re-issues a load. No test covers it. The new dataset builder won't run for anyone who installed GAIA normally. Its two new packages aren't listed in the package manifest, so the documented Also worth resolving before merge: the three new agent-loop guards fire only once per process rather than once per turn, so they stop protecting anything after the first trigger; and removing the step ceiling plus unpinning the context window are both model-behaviour changes that the project asks to be backed by an eval run. Real-world evidenceThe automated evidence stage failed before producing any evidence — This verdict therefore rests on static review alone. That matters more than usual for this PR, because the surfaces it changes are exactly the ones static review is weakest on: the shell tool's allowlist, the agent loop's stopping behaviour, and Lemonade model loading. Before merge I'd want to see a real 🔍 Technical details🔴 Critical1. Diff reverts commits already on main (whole-PR issue) The PR head genuinely carries the pre-fix content — I verified against the checked-out tree, not just the diff ( Reverted, with the commit each belongs to:
The 2.
if loaded_ctx >= expected_ctx: # TypeError: '>=' not supported between 'int' and 'NoneType'The Please add a 🟡 Important3. New factory subpackages missing from
4. The three new loop guards never reset between turns —
The new tests ( 5. 🔒 Shell allowlist widening needs maintainer sign-off — Three changes land together, and the reasoning for each is written up well; the question is whether the combination is the intended posture.
6. Unbounded steps + unpinned context, with no eval run —
Please attach a 7. Written-file guards resolve against
🟢 Minor
Strengths
|
…ow a declared sandbox Argument checking read tokens from shlex.split, which runs in POSIX mode and treats a backslash as an escape. A Windows-style absolute path tokenised with its separators removed, stopped looking like a path, and was never handed to PathValidator -- while the untokenised string is what reaches the shell, so the file was read. Absolute paths are now recovered from the raw command string and validated alongside the tokenised operands. Also adds GAIA_SHELL_SANDBOXED, off unless explicitly set, for callers whose blast radius really is disposable. It lifts the command-name allowlist and the operator ban only; path containment is unchanged and asserted in both modes.
|
Verdict: Request changes Two blocking regressions, both introduced by deleting tests first and then removing the code those tests enforced. 🔴 🔴 Email prescan regression. Two more concerns worth resolving before merge: 🟡 Action version downgrades. Nine workflow files downgrade 🔍 Technical detailsRemoteDisconnected sites left unguarded:
The deleted test file ( Email prescan regression:
lead = f"Here's your inbox pre-scan — {summary}. {coverage}."
if envelope.get("degraded"):
lead += " " + _mailbox_failure_caveat(envelope.get("mailbox_errors"))
return leadPost-patch: the The log-message rename in Action downgrade: |
|
Three of the defects here are also fixed in open PRs of mine, in the same files. Flagging the overlap so we don't land two fixes for the same thing — not to claim priority.
On Two fixes in mine that this PR does not appear to touch — I checked the diff for both:
Worth noting the two harnesses agree from different directions. Yours: refusals 55–76% → 6–10%, tasks 19/27 → 26/27. Mine: 14 tasks × 11 models × 3 GAIA commits, 322 runs scored by real pytest runs and probe scripts, where shell refusals were 17% of all tool calls across 58% of runs and the single largest source of wasted steps. Independent evidence for the same conclusion. Happy to close mine, rebase them onto this, or split them so each lands once — whichever keeps the tree simplest. Your call. |
|
Two things from running the flagship against a benchmark today — one is evidence for a change here, the other is a defect I think would fire false corrections in production. Your scratch-runner finding reproduces exactly. The comment here says the agent "could not run import sys
import pytest
sys.exit(pytest.main(["-q", "tests/"]))then a The unwritten-claims guard resolves against the wrong root.
So the agent writes Two open PRs widen the gap further: #3902 moves script execution to the resolved project root, and #3907 moves scratch writes to a dedicated scratch dir. After either lands, cwd is no longer where written files live even outside a sidecar. Suggested: resolve against 🔍 Technical detailsCall sites in the current head of this PR (from the PR files API —
Worth a test that the guard stays silent when the root cannot be resolved, given this PR's own history is a guard that never executed because its failure was swallowed. |
…timeout Adds GAIA_SHELL_SANDBOXED (off unless set) so a caller whose blast radius is genuinely disposable can declare it. Lifts the command-name allowlist and the operator ban only; path containment is unchanged and asserted in both modes. Measured against 28,064 recorded agent shell commands, 94% use an operator the default rules refuse, so a comparison against an agent run with permissions bypassed was not measuring what it claimed to. Also fixes a pre-existing path-validation bypass: argument checking read tokens from shlex.split in POSIX mode, which strips backslashes, so a Windows absolute path stopped looking like a path and was never validated -- while the untokenised string is what reaches the shell. Confirmed reading a canary outside the workspace with the allowlist fully enabled. And enforces the harness task timeout against the whole process tree; subprocess.run(timeout=) left one task running 92 minutes against a 900s limit, recorded as an agent failure.
|
Verdict: Approve with suggestions The core agent improvements — unlimited steps by default, idle-turn/repeat-call/missing-output guards, binary-sniff search, shell sandbox expansion — are well-designed and clearly motivated. Three findings below; the first two are worth resolving before merge. 🟡 Email: #3768 fix is silently reverted
If the fallback is now unreachable for degraded scans (perhaps the model's own framing always passes grounding for degraded envelopes), a one-sentence comment in 🟡 OAuth polling path: raw response body now in In If 🟢
🔍 Technical detailsEmail regression
OAuth raw body
Nudge flag reset
|
…-fixes # Conflicts: # docs/releases/v0.15.3.mdx # src/gaia/agents/base/agent.py # tests/unit/agents/test_agent_source_invariants.py # tests/unit/agents/test_loop_break_truthful.py # tests/unit/agents/test_verification_scope.py
…ed budget NO_STEP_LIMIT is 0, so 'steps_taken < steps_limit - 1' reads as 'steps_taken < -1' and is False forever. DEFAULT_MAX_STEPS is also 0, so on the default budget the guard added in #3750's follow-up could not fire at all. The eight other step comparisons in the same loop already carry the 'steps_limit == NO_STEP_LIMIT or ...' guard; this one was missed. Caught by the source invariant in test_default_max_steps.py, which scans for unguarded step comparisons -- the same silent-no-op shape as a guard whose helper was never imported.
…ed budget NO_STEP_LIMIT is 0, so 'steps_taken < steps_limit - 1' reads as 'steps_taken < -1' and is False forever. DEFAULT_MAX_STEPS is also 0, so on the default budget this guard could not fire at all. The eight other step comparisons in the same loop already carry the 'steps_limit == NO_STEP_LIMIT or ...' guard; this one was missed. Caught by the source invariant in test_default_max_steps.py, which scans for unguarded step comparisons -- the same silent-no-op shape as a guard whose helper was never imported.
|
Verdict: Request changes — two reversions need attention before merge. 🔴 OAuth error handling re-exposes raw provider bodies, reversing the #3875 fix.
The fix is to replicate 🔴 The #3768 degraded-scan fix is reversed without explanation.
🟡 Multiple 🔍 Technical details
# New — re-exposes raw body as fallback
error_description=err_payload.get("error_description", response.text[:300]),Old error_description=err_payload.get("error_description", ""),
# New — raw body in user-visible error
f"{resp.status_code}: {resp.text[:300]}. Check the client id "Old code extracted only
if envelope.get("degraded"):
lead += " " + _mailbox_failure_caveat(envelope.get("mailbox_errors"))
return leadwas the protection for #3768. The deleted
Action downgrade note: All |
|
Two blockers found while merging this into a test build. Both are verified against this PR's current head. 1. This PR would silently revert 8 fixes that are already on 2. The shell now refuses any relative path containing a slash. It fails with the sandbox switch both off and on. 🔍 Technical detailsReverts. For each commit on
Control: #3867 is also in the base and touches none of these files; this PR removes 0 of its 55 lines. Relative paths. Reproduced with |
An agent run could end with "Task completed with
<tool>. No further action needed." having written nothing at all — and nothing downstream could tell that apart from a real answer. The status wassuccess, every tool call returnedsuccess, and the file the user asked for did not exist. This fixes that and five more defects behind it, all found by running real tasks to completion and reading what the agent actually did rather than what it said.Measured across three-run batches on a 27-task suite: shell-tool refusals fell from 55–76% of calls to 6–10%, and tasks accomplished rose from 19/27 to 26/27.
The most useful finding is one that had been hiding a fix: the guard meant to catch a missing output file called a symbol that was never imported, and both call sites caught the resulting
NameErrorunder a blanketexcept Exception. It had never executed — so two earlier rounds of measurement were scoring a feature that did not exist. Its unit tests passed throughout, because they tested the regex and not the loop.🔍 What changed
NameError)search_filecould not find a file named exactly (ci.log,pyproject.toml).toml,.cfg,.tsand every extensionless file</think>left a reasoning model's deliberation in the answerpython -m, and its hint was untrueBoth output-guard handlers narrow from
except ExceptiontoOSError, so a programming error crashes instead of silently disabling a guard.The extension allowlists are replaced by a binary sniff (a NUL byte in the first 4 KB — the rule
grep -ruses) rather than extended. Any such list is wrong for the next language someone searches.Also included is the task-execution harness these were found with: the real agent loop run to completion, a fresh sandbox and process per attempt with
HOME/GAIA_HOMEredirected, correctness decided by a verifier's exit code, and judged quality kept in a separate column and never blended with it. Every verifier is proved two-sided — it must reject an untouched workspace and accept a hand-written oracle.Two scoped but not fixed here, because they are design decisions rather than bugs: #3967 (78 tools, 11 of them named
search_*; the model picked the content search zero times in 688 calls) and #3946 (python -m pytestbypasses the pytest grant policy).Test plan
python -m pytest tests/unit/agents tests/unit/factory tests/unit/test_shell_guardrails.py tests/unit/test_file_tools.py -q— 1517 passed locallypython util/lint.py --black --isort --flake8python -m pytest tests/unit/agents/test_loop_break_truthful.py tests/unit/agents/test_agent_source_invariants.py -qpython -m pytest tests/unit/agents/test_output_guards_in_loop.py -qpython -m pytest tests/unit/agents/test_python_console_scripts.py -q