fix(email-agent): auto-escalate search_messages to bodies for content questions - #3881
Conversation
… questions
A content question ("who signed this?", "what date was agreed?") came back
unanswerable even when the answer was in the mailbox, because
search_messages only fetched full bodies when include_bodies=True was
passed explicitly and a small local model did not reliably set it.
include_bodies now defaults to None (auto-decide): a pure Gmail operator
filter (from:/is:/label:/dates) stays metadata-only, while a query that
also carries a real search term escalates to full bodies automatically,
capped to an already-narrowed candidate set (SEARCH_AUTO_BODY_CAP=5) so a
broad content-shaped query still can't reproduce the #2763 overflow.
Explicit include_bodies=True/False still override the heuristic, uncapped.
pre_scan_inbox's docstrings now state plainly that no message body is ever
read on that surface, and point at get_message/search_messages instead.
Wiring an actual body-reading path into pre-scan is #2968, not done here.
Closes #3773
Skill audit
✅ All audited skills cleared the tier they claim. Per-finding detail is withheld here on purpose. Read it in the Security > Code scanning tab, or download the |
|
Verdict: Request changes — small, targeted; the core idea is sound. Search now decides for itself when to read full message bodies, which is the right fix for "the model never opted in". The selector, the cap, and the overrides are well tested, and the CI evidence run shows the escalation working end to end against a hermetic mailbox. Two things to fix before merge:
Also: this changes the instructions the model reads for a tool, which the project requires an agent eval to cover. The PR already calls that out as required-before-merge and not yet run — that still needs to happen (it needs a machine with local inference). Real-world evidenceAn evidence bundle ran on a no-inference runner and exercised the changed tool directly, with a planted body fact: The tool schema the model receives still advertises a plain optional boolean ( Two gaps the bundle states itself: the rendered Agent UI turn and real model tool-choice are pending the strix-halo lane (no local inference on this runner), and the retry case it ran matched nothing, so — in its own words — "it does not exercise a retried query that also returns hits." That is precisely the gap in the first finding above, so that finding rests on static review plus the confirmed classifier behaviour, not on a live failure. 🔍 Technical details🟡 Auto-escalation is structurally impossible on the operator-retry path ( The retry only fires when the original query carried no operator ( So every #2114-style search (literal phrase → 0 hits → operator retry → hits) returns metadata-only, which is the #3773 symptom. The inline rationale ("checking the original bare phrase would misread an already-operator-only retried query") is inverted: a retried query is by construction only ever produced from a bare-phrase query. Worth a test in 🟡 The wrapper calls Relatedly, 🟡 Agent eval not run for an LLM-affecting change (
🟢 Nit — Strengths
|
Summary
A content question like "who signed this?" or "what date was agreed?" came back unanswerable even when the answer was sitting in the mailbox.
search_messagesonly fetched full message bodies wheninclude_bodies=Truewas passed explicitly, and a small local model did not reliably set it — most content questions got a metadata-only answer or a flat refusal.include_bodiesnow defaults to auto-deciding from the search query's shape — a syntactic check, not an understanding of the question: a pure Gmail filter (from:/is:/label:/date operators) still returns metadata only, while a query that also carries a bare search term escalates to full bodies automatically, capped to an already-narrowed candidate set (5 messages) so a broad content-shaped query still can't overflow the model's context window the way #2763 fixed. An explicitinclude_bodies=True/Falsestill overrides this and is never capped.pre_scan_inbox's docstrings now say plainly that no message body is ever read on that surface (unchanged behavior, documented for the first time, with the caller pointed atget_message/search_messagesinstead) — see the design notes below for why that surface itself was left alone.Closes #3773
Design notes and known limitation
_query_has_free_textis a syntactic discriminator over the Gmail query the model already has to write — it keys on query shape (does it carry a term beyond known structural operators), not on any understanding of what's being asked.SEARCH_AUTO_BODY_CAP = 5bounds the auto-escalation to an already-narrowed candidate set, per the acceptance criteria.pre_scan_inboxstill never wires aclassifier=intotriage_inbox_impl, so it still never reads a body on that surface. Wiring an actual bounded body-reading path there is fix(email): CLI triage skips categories; UI path works fine #2968 — deliberately out of scope here, per the adjacent-issue note in the original task. This PR only makes that existing behavior explicit in the docstrings the model and future maintainers read.body_truncated/body_chars_dropped) already existed on every full-body formatter since fix(email): give list_inbox/search_messages a combined envelope budget #2546 and needed no new code. Verified empirically that it flows through the new auto-escalation path unchanged (it reuses the same_format_messages_within_budget→_format_message_for_llmpath as the explicitinclude_bodies=Truecase): a message with a body 500 chars over the limit reportsbody_truncated: True, body_chars_dropped: 500when auto-escalated.What is demonstrated vs. not
Test plan
test_search_messages_body_escalation_3773.py(19 tests) pins the selector: which queries escalate, the cap boundary (at/overSEARCH_AUTO_BODY_CAP), that a pure-filter/counting query escalates none of its results regardless of hit count, and that explicitinclude_bodies=True/Falsealways overrides the heuristic uncapped.pytest hub/agents/email -q(deselecting one pre-existing Lemonade-dependent SLM test and one pre-existing packaging test needing the[publish]extra — both unrelated to this change, confirmed failing the same way on an unmodified tree) — 2089 passed.pytest tests/unit -k email -q(deselecting one pre-existinggaia-agent-emailconsole-script PATH check, an environment/install artifact unrelated to this change) — 1524 passed.python util/lint.py --all— black/isort clean on touched files; the one mypy error reported is pre-existing and in an unrelated file (src/gaia/factory/harvest/scan.py).gaia eval agent --category tool_selectionvs. baselinetests/fixtures/eval_baselines/gemma-4-e4b-d71cd914/scorecard_tool_selection.json— required before merge. Not run this session: a Lemonade backend is reachable, but the eval judge client is currently broken on its own default model (claude-opus-5returns a['thinking','text']content block ordering thatsrc/gaia/eval/claude.pydoesn't handle,AttributeErroron the first reply — tracked as fix(eval): the judge crashes on its own default model — content[0] is a thinking block #3884, fix in progress). This is not "skipped" and not "passing" — it is the one thing in this PR that has not been checked.