Skip to content

Use model token IDs for guidance markers - #1093

Merged
Baiju Meswani (baijumeswani) merged 4 commits into
mainfrom
baijumeswani/literal-marker-guidance
Sep 11, 2026
Merged

Baiju Meswani (baijumeswani) merged 4 commits into
mainfrom
baijumeswani/literal-marker-guidance

Conversation

@baijumeswani

@baijumeswani Baiju Meswani (baijumeswani) commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator

What this fixes

Foundry Local builds an llguidance grammar when a request requires a tool call.

Previously, Foundry inserted marker text directly into the grammar:

toolcall: <tool_call> functioncall </tool_call>

In llguidance, <tool_call> means a named tokenizer special token. This fails for models that use the same marker as
an ordinary token or ordinary text.

What changed

Foundry now uses the boundary token IDs published by ORT GenAI:

  • BOT — beginning of tool call
  • EOT — end of tool call
  • BOR — beginning of reasoning
  • EOR — end of reasoning

When a published ID matches the configured marker, Foundry uses llguidance's exact numeric token syntax:

toolcall: <[248058]> functioncall <[248059]>

This works whether the tokenizer marks the token as special or ordinary.

If an ID is unavailable, invalid, or does not match an overridden marker, Foundry uses an escaped quoted string:

toolcall: "<tool_call>" functioncall "</tool_call>"

The quoted form allows llguidance to tokenize the marker as ordinary text, including markers represented by multiple
tokens.

Tool-call and reasoning start/end boundaries are resolved independently. The implementation contains no model names
or hardcoded token IDs.

Why not mark every marker as special?

Changing tokenizer metadata affects more than grammar construction. ORT GenAI normally decodes with
skip_special_tokens=true, so changing an ordinary marker to special: true can remove it from decoded output.

Using numeric IDs preserves the model package's tokenizer behavior while still constraining the exact published
token.

Required tool-call failures

If Foundry generates guidance for a required tool call and the runtime cannot apply it, the request now fails clearly
instead of continuing with unconstrained generation.

Invalid caller-provided guidance remains an invalid-argument error on both Generator and Engine backends.

Example

Request:

Use the multiply tool to calculate 17 times 23. Do not answer directly.

Result:

{
  "finish_reason": "tool_calls",
  "tool_calls": [
    {
      "type": "function",
      "function": {
        "name": "multiply",
        "arguments": "{\"a\":17,\"b\":23}"
      }
    }
  ]
}

After the tool returns 391, the next turn answers 17 × 23 = 391.

@vercel

vercel Bot commented Sep 9, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
foundry-local Ready Ready Preview Sep 11, 2026 1:21am UTC

Request Review

Base automatically changed from baijumeswani/tool-transcript to main September 10, 2026 21:32
Render model-published tool and reasoning boundary IDs with llguidance numeric token syntax. Fall back to escaped string literals when a boundary ID is unavailable or does not match the configured marker.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e1b3adfd-19d8-4423-8a48-328526fda78f
@baijumeswani Baiju Meswani (baijumeswani) changed the title Render non-special markers as guidance literals Use exact token IDs for guidance markers Sep 10, 2026
@baijumeswani
Baiju Meswani (baijumeswani) force-pushed the baijumeswani/literal-marker-guidance branch from 00f5a98 to b07008d Compare September 10, 2026 23:13
Copilot AI balanced review requested due to automatic review settings September 10, 2026 23:13
@baijumeswani Baiju Meswani (baijumeswani) changed the title Use exact token IDs for guidance markers Use model token IDs for guidance markers Sep 10, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The grammar helper can render invalid numeric-token syntax for negative or sentinel IDs.

Pull request overview

Updates chat guidance to use exact published token IDs, with escaped-literal fallback and safer required-tool handling across Generator and Engine backends.

Changes:

  • Added numeric-token grammar rendering and literal escaping.
  • Propagated boundary IDs and reasoning state through chat flows.
  • Expanded grammar, planning, and Engine tests.
File summaries
File Description
sdk_v2/cpp/test/internal_api/toolcalling/grammar_test.cc Tests numeric-token and literal rendering.
sdk_v2/cpp/test/internal_api/chat/search_options_test.cc Tests guidance planning and reasoning state.
sdk_v2/cpp/test/internal_api/chat/dynamic_engine_chat_test.cc Updates Engine turn tests.
sdk_v2/cpp/src/inferencing/generative/toolcalling/tool_call_context.h Adds boundary token ID fields.
sdk_v2/cpp/src/inferencing/generative/toolcalling/grammar.h Declares grammar rendering APIs.
sdk_v2/cpp/src/inferencing/generative/toolcalling/grammar.cc Implements token rendering and escaping.
sdk_v2/cpp/src/inferencing/generative/genai_model_instance.h Declares tokenizer encoding support.
sdk_v2/cpp/src/inferencing/generative/genai_model_instance.cc Implements tokenizer encoding support.
sdk_v2/cpp/src/inferencing/generative/chat/search_options.h Defines guidance planning interfaces.
sdk_v2/cpp/src/inferencing/generative/chat/search_options.cc Handles guidance planning and errors.
sdk_v2/cpp/src/inferencing/generative/chat/onnx_engine_chat_stream.cc Propagates reasoning state.
sdk_v2/cpp/src/inferencing/generative/chat/onnx_chat_generator.cc Resolves markers and applies guidance.
sdk_v2/cpp/src/inferencing/generative/chat/onnx_chat_engine.h Extends Engine guidance interfaces.
sdk_v2/cpp/src/inferencing/generative/chat/onnx_chat_engine.cc Integrates Engine guidance handling.
sdk_v2/cpp/src/inferencing/generative/chat/chat_session.cc Resolves published boundary IDs.
Review details

Suppressed comments (1)

sdk_v2/cpp/src/inferencing/generative/toolcalling/grammar.cc:59

  • This treats any engaged ID as valid, so a sentinel or malformed negative ID produces <[-1]>, which is not valid llguidance numeric-token syntax. ResolveMarkerTokenId filters negatives on the current ChatSession path, but this public helper can also be called with a directly populated ToolCallContext; only nonnegative IDs should use exact-token rendering, with invalid IDs falling back to the escaped marker literal.
  if (token_id.has_value()) {
    return "<[" + std::to_string(*token_id) + "]>";
  • Files reviewed: 15/15 changed files
  • Comments generated: 0
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread sdk_v2/cpp/src/inferencing/generative/toolcalling/grammar.cc
Comment thread sdk_v2/cpp/test/internal_api/toolcalling/grammar_test.cc
Fall back to quoted marker text for negative token IDs and add explicit Phi-4 mini reasoning coverage when no boundary IDs are published.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e1b3adfd-19d8-4423-8a48-328526fda78f
Use an ordinary escaped string instead of a raw literal whose closing quote is ambiguous to the Windows compiler.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e1b3adfd-19d8-4423-8a48-328526fda78f
Use explicit ordinary C++ strings for quote, slash, and control-character expectations so the assertions parse consistently under MSVC.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e1b3adfd-19d8-4423-8a48-328526fda78f
@baijumeswani
Baiju Meswani (baijumeswani) merged commit b3f65c4 into main Sep 11, 2026
60 checks passed
@baijumeswani
Baiju Meswani (baijumeswani) deleted the baijumeswani/literal-marker-guidance branch September 11, 2026 02:46

This branch was successfully deployed

1 active deployment
Preview — 1746b968 Deployed Sep 11, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants