Use model token IDs for guidance markers - #1093
Merged
Baiju Meswani (baijumeswani) merged 4 commits intoSep 11, 2026
Merged
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Baiju Meswani (baijumeswani)
force-pushed
the
baijumeswani/tool-transcript
branch
from
September 10, 2026 18:33
dc35ef9 to
03c17c3
Compare
Render model-published tool and reasoning boundary IDs with llguidance numeric token syntax. Fall back to escaped string literals when a boundary ID is unavailable or does not match the configured marker. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e1b3adfd-19d8-4423-8a48-328526fda78f
Baiju Meswani (baijumeswani)
force-pushed
the
baijumeswani/literal-marker-guidance
branch
from
September 10, 2026 23:13
00f5a98 to
b07008d
Compare
Copilot started reviewing on behalf of
Baiju Meswani (baijumeswani)
September 10, 2026 23:14
View session
Contributor
There was a problem hiding this comment.
🔵 Needs a closer look
The grammar helper can render invalid numeric-token syntax for negative or sentinel IDs.
Pull request overview
Updates chat guidance to use exact published token IDs, with escaped-literal fallback and safer required-tool handling across Generator and Engine backends.
Changes:
- Added numeric-token grammar rendering and literal escaping.
- Propagated boundary IDs and reasoning state through chat flows.
- Expanded grammar, planning, and Engine tests.
File summaries
| File | Description |
|---|---|
sdk_v2/cpp/test/internal_api/toolcalling/grammar_test.cc |
Tests numeric-token and literal rendering. |
sdk_v2/cpp/test/internal_api/chat/search_options_test.cc |
Tests guidance planning and reasoning state. |
sdk_v2/cpp/test/internal_api/chat/dynamic_engine_chat_test.cc |
Updates Engine turn tests. |
sdk_v2/cpp/src/inferencing/generative/toolcalling/tool_call_context.h |
Adds boundary token ID fields. |
sdk_v2/cpp/src/inferencing/generative/toolcalling/grammar.h |
Declares grammar rendering APIs. |
sdk_v2/cpp/src/inferencing/generative/toolcalling/grammar.cc |
Implements token rendering and escaping. |
sdk_v2/cpp/src/inferencing/generative/genai_model_instance.h |
Declares tokenizer encoding support. |
sdk_v2/cpp/src/inferencing/generative/genai_model_instance.cc |
Implements tokenizer encoding support. |
sdk_v2/cpp/src/inferencing/generative/chat/search_options.h |
Defines guidance planning interfaces. |
sdk_v2/cpp/src/inferencing/generative/chat/search_options.cc |
Handles guidance planning and errors. |
sdk_v2/cpp/src/inferencing/generative/chat/onnx_engine_chat_stream.cc |
Propagates reasoning state. |
sdk_v2/cpp/src/inferencing/generative/chat/onnx_chat_generator.cc |
Resolves markers and applies guidance. |
sdk_v2/cpp/src/inferencing/generative/chat/onnx_chat_engine.h |
Extends Engine guidance interfaces. |
sdk_v2/cpp/src/inferencing/generative/chat/onnx_chat_engine.cc |
Integrates Engine guidance handling. |
sdk_v2/cpp/src/inferencing/generative/chat/chat_session.cc |
Resolves published boundary IDs. |
Review details
Suppressed comments (1)
sdk_v2/cpp/src/inferencing/generative/toolcalling/grammar.cc:59
- This treats any engaged ID as valid, so a sentinel or malformed negative ID produces
<[-1]>, which is not valid llguidance numeric-token syntax.ResolveMarkerTokenIdfilters negatives on the current ChatSession path, but this public helper can also be called with a directly populatedToolCallContext; only nonnegative IDs should use exact-token rendering, with invalid IDs falling back to the escaped marker literal.
if (token_id.has_value()) {
return "<[" + std::to_string(*token_id) + "]>";
- Files reviewed: 15/15 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Fall back to quoted marker text for negative token IDs and add explicit Phi-4 mini reasoning coverage when no boundary IDs are published. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e1b3adfd-19d8-4423-8a48-328526fda78f
kunal-vaishnavi
previously approved these changes
Sep 10, 2026
Baiju Meswani (baijumeswani)
enabled auto-merge (squash)
September 10, 2026 23:48
Use an ordinary escaped string instead of a raw literal whose closing quote is ambiguous to the Windows compiler. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e1b3adfd-19d8-4423-8a48-328526fda78f
kunal-vaishnavi
previously approved these changes
Sep 11, 2026
Use explicit ordinary C++ strings for quote, slash, and control-character expectations so the assertions parse consistently under MSVC. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e1b3adfd-19d8-4423-8a48-328526fda78f
kunal-vaishnavi
approved these changes
Sep 11, 2026
Baiju Meswani (baijumeswani)
deleted the
baijumeswani/literal-marker-guidance
branch
September 11, 2026 02:46
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this fixes
Foundry Local builds an llguidance grammar when a request requires a tool call.
Previously, Foundry inserted marker text directly into the grammar:
In llguidance,
<tool_call>means a named tokenizer special token. This fails for models that use the same marker asan ordinary token or ordinary text.
What changed
Foundry now uses the boundary token IDs published by ORT GenAI:
When a published ID matches the configured marker, Foundry uses llguidance's exact numeric token syntax:
This works whether the tokenizer marks the token as special or ordinary.
If an ID is unavailable, invalid, or does not match an overridden marker, Foundry uses an escaped quoted string:
The quoted form allows llguidance to tokenize the marker as ordinary text, including markers represented by multiple
tokens.
Tool-call and reasoning start/end boundaries are resolved independently. The implementation contains no model names
or hardcoded token IDs.
Why not mark every marker as special?
Changing tokenizer metadata affects more than grammar construction. ORT GenAI normally decodes with
skip_special_tokens=true, so changing an ordinary marker tospecial: truecan remove it from decoded output.Using numeric IDs preserves the model package's tokenizer behavior while still constraining the exact published
token.
Required tool-call failures
If Foundry generates guidance for a required tool call and the runtime cannot apply it, the request now fails clearly
instead of continuing with unconstrained generation.
Invalid caller-provided guidance remains an invalid-argument error on both Generator and Engine backends.
Example
Request:
Result:
{ "finish_reason": "tool_calls", "tool_calls": [ { "type": "function", "function": { "name": "multiply", "arguments": "{\"a\":17,\"b\":23}" } } ] }After the tool returns
391, the next turn answers17 × 23 = 391.