Skip to content

fix(email-agent): auto-escalate search_messages to bodies for content questions - #3881

Merged
itomek merged 1 commit into
mainfrom
issue-3773
Sep 15, 2026
Merged

itomek merged 1 commit into
mainfrom
issue-3773

Conversation

@itomek

@itomek itomek commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A content question like "who signed this?" or "what date was agreed?" came back unanswerable even when the answer was sitting in the mailbox. search_messages only fetched full message bodies when include_bodies=True was passed explicitly, and a small local model did not reliably set it — most content questions got a metadata-only answer or a flat refusal.

include_bodies now defaults to auto-deciding from the search query's shape — a syntactic check, not an understanding of the question: a pure Gmail filter (from:/is:/label:/date operators) still returns metadata only, while a query that also carries a bare search term escalates to full bodies automatically, capped to an already-narrowed candidate set (5 messages) so a broad content-shaped query still can't overflow the model's context window the way #2763 fixed. An explicit include_bodies=True/False still overrides this and is never capped.

pre_scan_inbox's docstrings now say plainly that no message body is ever read on that surface (unchanged behavior, documented for the first time, with the caller pointed at get_message/search_messages instead) — see the design notes below for why that surface itself was left alone.

Closes #3773

Design notes and known limitation
  • The issue's own proposed mechanism was a natural-language "is this a content question vs. a counting question" classifier. That was rejected: it reintroduces the exact small-model-unreliability problem fix(email-agent): search_messages returns no answer at all for a long-bodied sender — context overflow #2763 already hit, just moved from a boolean flag to fuzzier free-text classification. Instead, _query_has_free_text is a syntactic discriminator over the Gmail query the model already has to write — it keys on query shape (does it carry a term beyond known structural operators), not on any understanding of what's being asked.
  • Known boundary of the heuristic, stated explicitly: because it is syntactic, a bare-keyword query that is really a counting question ("how many emails mention refund") will escalate to bodies unnecessarily, and a genuine content question phrased purely in operators ("from:acme is:unread", hoping to read the one match) will not. The heuristic optimizes for the common case (a content question naturally includes the words being asked about; a counting question naturally doesn't) rather than claiming semantic understanding.
  • SEARCH_AUTO_BODY_CAP = 5 bounds the auto-escalation to an already-narrowed candidate set, per the acceptance criteria.
  • pre_scan_inbox still never wires a classifier= into triage_inbox_impl, so it still never reads a body on that surface. Wiring an actual bounded body-reading path there is fix(email): CLI triage skips categories; UI path works fine #2968 — deliberately out of scope here, per the adjacent-issue note in the original task. This PR only makes that existing behavior explicit in the docstrings the model and future maintainers read.
  • Truncation surfacing (acceptance criterion 5: body_truncated/body_chars_dropped) already existed on every full-body formatter since fix(email): give list_inbox/search_messages a combined envelope budget #2546 and needed no new code. Verified empirically that it flows through the new auto-escalation path unchanged (it reuses the same _format_messages_within_budget → _format_message_for_llm path as the explicit include_bodies=True case): a message with a body 500 chars over the limit reports body_truncated: True, body_chars_dropped: 500 when auto-escalated.

What is demonstrated vs. not

  • Demonstrated (hermetic, in this PR): the escalation selector — which queries escalate, the cap boundary, and specifically that a metadata-answerable (pure-filter/counting) request escalates none of its results regardless of hit count. This is fully covered by the new unit tests below.
  • Not demonstrated (requires live inference): whether a real local model, given this change, actually produces a better answer to a content question end-to-end. That is a model-behavior question the eval measures, not something the selector's unit tests can prove.

Test plan

  • New unit test file test_search_messages_body_escalation_3773.py (19 tests) pins the selector: which queries escalate, the cap boundary (at/over SEARCH_AUTO_BODY_CAP), that a pure-filter/counting query escalates none of its results regardless of hit count, and that explicit include_bodies=True/False always overrides the heuristic uncapped.
  • pytest hub/agents/email -q (deselecting one pre-existing Lemonade-dependent SLM test and one pre-existing packaging test needing the [publish] extra — both unrelated to this change, confirmed failing the same way on an unmodified tree) — 2089 passed.
  • pytest tests/unit -k email -q (deselecting one pre-existing gaia-agent-email console-script PATH check, an environment/install artifact unrelated to this change) — 1524 passed.
  • python util/lint.py --all — black/isort clean on touched files; the one mypy error reported is pre-existing and in an unrelated file (src/gaia/factory/harvest/scan.py).
  • gaia eval agent --category tool_selection vs. baseline tests/fixtures/eval_baselines/gemma-4-e4b-d71cd914/scorecard_tool_selection.json — required before merge. Not run this session: a Lemonade backend is reachable, but the eval judge client is currently broken on its own default model (claude-opus-5 returns a ['thinking','text'] content block ordering that src/gaia/eval/claude.py doesn't handle, AttributeError on the first reply — tracked as fix(eval): the judge crashes on its own default model — content[0] is a thinking block #3884, fix in progress). This is not "skipped" and not "passing" — it is the one thing in this PR that has not been checked.

… questions

A content question ("who signed this?", "what date was agreed?") came back
unanswerable even when the answer was in the mailbox, because
search_messages only fetched full bodies when include_bodies=True was
passed explicitly and a small local model did not reliably set it.

include_bodies now defaults to None (auto-decide): a pure Gmail operator
filter (from:/is:/label:/dates) stays metadata-only, while a query that
also carries a real search term escalates to full bodies automatically,
capped to an already-narrowed candidate set (SEARCH_AUTO_BODY_CAP=5) so a
broad content-shaped query still can't reproduce the #2763 overflow.
Explicit include_bodies=True/False still override the heuristic, uncapped.

pre_scan_inbox's docstrings now state plainly that no message body is ever
read on that surface, and point at get_message/search_messages instead.
Wiring an actual body-reading path into pre-scan is #2968, not done here.

Closes #3773
@github-actions github-actions Bot added the agent::email Email agent changes label Sep 15, 2026
@itomek
itomek marked this pull request as ready for review September 15, 2026 04:59
@github-actions

Copy link
Copy Markdown
Contributor

Skill audit

Skill Verdict Claimed tier Cleared tiers Findings Rules
hub/agents/email/npm ✅ ALLOW experimental experimental, community 2 medium, 1 info code.suppression, supply.undeclared_dependency

✅ All audited skills cleared the tier they claim.

Per-finding detail is withheld here on purpose. Read it in the Security > Code scanning tab, or download the skill-audit-reports artifact from this run. Offending source text is withheld from CI everywhere — reproduce it locally with gaia skill audit <dir> --show-snippets.

@github-actions

Copy link
Copy Markdown
Contributor

Verdict: Request changes — small, targeted; the core idea is sound.

Search now decides for itself when to read full message bodies, which is the right fix for "the model never opted in". The selector, the cap, and the overrides are well tested, and the CI evidence run shows the escalation working end to end against a hermetic mailbox.

Two things to fix before merge:

  • The fix doesn't reach the case it was written for. When a plain-words query finds nothing literally, search retries it as a sender/subject query — and on that retry path the new auto-escalation can never fire, so a question like "what did the Netflix email say?" still comes back without the message text. The retry only ever runs for plain-words queries, so this suppresses exactly the searches the change is meant to help. Decide from the words the user searched for, not from the rewritten retry query, and add a test where the retry actually returns hits.
  • With two mailboxes connected, the "handful of messages" guard is per mailbox, not per answer. Two connected accounts can each escalate their own small batch, so a single search can return roughly twice the body text the cap implies — and the result is labelled as auto-escalated even when only one mailbox's entries actually carry bodies, which invites the model to say "nothing in the other mailbox mentions this" after only reading subjects there.

Also: this changes the instructions the model reads for a tool, which the project requires an agent eval to cover. The PR already calls that out as required-before-merge and not yet run — that still needs to happen (it needs a machine with local inference).

Real-world evidence

An evidence bundle ran on a no-inference runner and exercised the changed tool directly, with a planted body fact:

=== content-shaped query, 2 hits (<= cap) — the #3773 fix
$ search_messages(query='contract renewal', max_results=25)
{ "count": 2, "include_bodies_auto": true,
  "messages[0]_body": "...countersigned by Zephyr Okonkwo on 2026-03-14, ref violet-otter-92..." }

=== pure filter query, 1 hit — counting/listing stays metadata-only
$ search_messages(query='from:vendor', max_results=25)
{ "count": 1, "include_bodies_auto": false, "messages[0]_body": null }

=== content-shaped query, 6 hits (> cap) — no overflow
{ "count": 6, "include_bodies_auto": false, "messages[0]_body": null }

=== explicit include_bodies=True on a pure filter, 8 hits — override, uncapped
{ "count": 8, "include_bodies_auto": false, "messages[0]_body": "...Zephyr Okonkwo..." }

The tool schema the model receives still advertises a plain optional boolean ('include_bodies': {'type': 'boolean', 'required': False}), and the email sidecar booted with /v1/email/health → 200, /v1/email/prescan → 503 with the connect-a-mailbox message.

Two gaps the bundle states itself: the rendered Agent UI turn and real model tool-choice are pending the strix-halo lane (no local inference on this runner), and the retry case it ran matched nothing, so — in its own words — "it does not exercise a retried query that also returns hits." That is precisely the gap in the first finding above, so that finding rests on static review plus the confirmed classifier behaviour, not on a live failure.

🔍 Technical details

🟡 Auto-escalation is structurally impossible on the operator-retry path (read_tools.py:1030-1038)

The retry only fires when the original query carried no operator (read_tools.py:1017), and operatorize_query rewrites it to from:(phrase) OR subject:(phrase). Deciding from effective_query therefore always yields False — verified:

True  'who signed the NDA'
False 'from:(who signed the NDA) OR subject:(who signed the NDA)'

So every #2114-style search (literal phrase → 0 hits → operator retry → hits) returns metadata-only, which is the #3773 symptom. The inline rationale ("checking the original bare phrase would misread an already-operator-only retried query") is inverted: a retried query is by construction only ever produced from a bare-phrase query.

            # #3773: the caller's own query is what signals content intent.
            # The #2114 operator retry is synthesized FROM a bare-phrase
            # query, so deciding on it alone would never escalate.
            effective_query = retried_query if retried_query is not None else query
            resolved_include_bodies = (
                _query_has_free_text(query) or _query_has_free_text(effective_query)
            ) and len(stubs) <= SEARCH_AUTO_BODY_CAP

Worth a test in test_search_messages_body_escalation_3773.py where the literal query misses and the retried query returns ≤ SEARCH_AUTO_BODY_CAP hits.

🟡 SEARCH_AUTO_BODY_CAP and the envelope budget are both per-backend (read_tools.py:3036-3053)

The wrapper calls search_messages_impl once per connected mailbox, each with budget_tokens=None (full device profile budget) and its own ≤5 escalation. With Google + Microsoft connected, a single search_messages can auto-escalate up to 10 full bodies against ~2× the intended context budget — reachable now without the model opting in, whereas pre-#3773 it required explicit include_bodies=True. Consider capping across the merge (e.g. escalate only if the merged candidate set is ≤ cap, or split the budget by len(backends) as per_backend already does for max_results).

Relatedly, include_bodies_auto = include_bodies_auto or bool(...) makes the flag true when any backend escalated, while the merged messages list can mix body-bearing and metadata-only entries. The docstring tells the model include_bodies_auto means "THIS call escalated to full bodies", so a mixed result can be read as "all hits were body-searched".

🟡 Agent eval not run for an LLM-affecting change (read_tools.py:2956-3025)

search_messages' model-facing docstring and pre_scan_inbox's output guidance both changed — CLAUDE.md requires gaia eval agent against the committed baseline for tool-docstring changes. The PR states this explicitly as required-before-merge and unrun (no Lemonade backend reachable), which is honest and correct for this lane; flagging so it isn't lost at merge time. Worth watching for the model now passing include_bodies=False explicitly out of habit from the old docstring, which silently disables the new default.

🟢 Nit — msgs is unused in most new tests (test_search_messages_body_escalation_3773.py:496, :509, :519, :530, :544, :556) — gmail, _ = _build_inbox(...) reads slightly cleaner where the list isn't asserted on.

Strengths

  • The syntactic selector is the right call over an LLM-based "is this a content question" classifier, and the PR says so with the reasoning — the fix(email-agent): search_messages returns no answer at all for a long-bodied sender — context overflow #2763 lesson applied rather than restated.
  • _OPERATOR_VALUE_RE derives from the canonical _GMAIL_OPERATORS tuple, so it can't drift from has_gmail_operator; the classifier behaves correctly on negated operators, parenthesized values, and quoted phrases.
  • Tests pin the selector (which queries escalate, the cap boundary at and above SEARCH_AUTO_BODY_CAP, both explicit overrides) at both the impl and registered-@tool layers, not just a return value.
  • Optional[bool] retyping was checked against the live tool registry rather than assumed — the schema still reads boolean/optional.
  • pre_scan_inbox's "no body is ever read" guarantee being documented on both the impl and the model-facing docstring is a genuine clarity win, and fix(email): CLI triage skips categories; UI path works fine #2968 is correctly left out of scope. No SPEC/SKILL/README drift: neither doc describes include_bodies, and both CHANGELOGs are updated.

@itomek
itomek added this pull request to the merge queue Sep 15, 2026
Merged via the queue into main with commit 09e4291 Sep 15, 2026
52 of 54 checks passed
@itomek
itomek deleted the issue-3773 branch September 15, 2026 14:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agent::email Email agent changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(email): triage and search answer from metadata and snippets, never message bodies

1 participant