test(eval): pin credential and MCP launch contracts - #3858
Conversation
Verdict: ApproveThis adds regression tests for the eval launcher's credential handling — the thing that silently broke authentication for subscription-only users — plus a smoke test that actually starts the configured MCP server. Tests only, no production code touched. Nothing blocking. The one gap worth knowing about: the smoke test launches the MCP server with the interpreter pytest is already running under, not the one the eval config actually names. So if that interpreter goes missing from the machine's search path, the eval breaks at spawn time and this test still passes green. The import chain it set out to check is covered correctly; it's the launch command itself that isn't. Real-world evidenceN/A — tests-only change, no user-facing surface. No 🔍 Technical detailsIssues🟢 Smoke test ignores the config's The test reads (requires 🟢 File placement diverges from the existing eval test package (
🟢 The It doubles the matrix to pin an unstated contract — that OAuth presence must not alter the command. Worth one short comment saying so, otherwise a future reader is likely to delete it as redundant. Notes (not findings)
Strengths
|
The test spawned the launcher with pytest's own interpreter while the eval spawns whatever the config's `command` resolves to on PATH. If that interpreter went missing the eval would break at spawn and this test would still pass — the one failure it exists to catch. It now uses the configured command and fails with a named reason when it is not resolvable.
|
Verdict: Approve Two new tests that pin the credential-selection and MCP-launch contracts for The logic is correct and well-matched to the implementation:
No correctness bugs, no security issues, no missing tests for new logic (this PR is the tests). Clean to merge. |
Agent evals now have regression coverage for the launch configuration that previously broke authentication: API-key/OAuth combinations, bare-mode selection, MCP configuration, and working directory. A real launcher smoke test also catches missing server imports before an eval starts.
Fixes #3569.
Test plan:
--help.These checks do not make live authenticated requests or run model evaluations.