Skip to content

FEAT: Generated transport conformance tests should cover new reject/dlq/requeue paths - #4297

Merged
iancooper merged 181 commits into
masterfrom
feature/4240-universal-transport-conformance-tests
Sep 19, 2026
Merged

iancooper merged 181 commits into
masterfrom
feature/4240-universal-transport-conformance-tests

Conversation

@iancooper

@iancooper iancooper commented Sep 1, 2026 •

Copy link
Copy Markdown
Member

Generated transport conformance tests - should cover new reject/dlq/requeue paths

Contributes to #4240, which stays open deliberately.

All 46 Deferred cells in the conformance ledger, and every generated Skip marker, resolve to
#4240. Closing it on merge would leave each of those pointing at a closed issue, and
LedgerSkipCrossCheckAudit never queries the tracker (ADR 0067 step 7), so nothing in CI would
ever notice - a silent skip with extra steps, in a PR whose central claim is that there are none.
#4240 therefore remains the umbrella tracker for the deferred behaviours, and carries a comment
recording what this PR completed.

While generating the transport conformance tests, I added new functionality around reject/dlq/requeue. This functionality can now be universal; where the message-oriented middleware doesn't support it, we can fall back to a Brighter substitute.

This replaces the four legacy per-transport gate keys, which controlled generation of functionality for these features, dependent on native support, with a generated canonical conformance suite that runs
against every gateway.

As some gateways have not yet implemented Brighter support for missing features, we also added a checked-in conformance ledger (specs/0036-universal-transport-conformance-tests/conformance-status.md) recording per-configuration conformance for 11 canonical behaviors (FR-2, 4, 5, 6, 7, 8, 9, 15, 16, 17, 22).

Every canonical test's Skip is driven by this ledger, so a behavior is either proven against a real broker or carries a signed-off Deferred -> #4240 marker. No silent skips.

Status — 62 of 62 tasks; phases 0–6 complete

All 24 wired configuration rows are resolved — zero Unknown cells, across all twelve targeted
transports:

Configuration Result
AWS V3 x4 + V4 x4 11 Pass, except SqsFifo FR-9 Deferred (SQS FIFO rejects per-message delay) and Sns* FR-9 Fixed
Kafka Classic / Consumer / PartitionKey 11/11 Pass
PostgresSQL 11/11 Pass (native visibility-timeout redelivery)
MSSQL, Redis 10 Pass + FR-16 Deferred (destructive read)
RMQ.Async Classic / Quorum 10 Pass + FR-5 Deferred (invalid-channel routing)
RMQ.Sync 10 Fixed + FR-5 Deferred, mirroring RMQ.Async
MQTT 10 Fixed + FR-16 Deferred
RocketMQ 9 Fixed + FR-2 / FR-15 Deferred (upstream no-op Requeue)
GCP Pull / PullOrdering 4 Pass + FR-9 Fixed + 6 Deferred (emulator-only verification)
GCP Stream / StreamOrdering all 11 Deferred — streaming pull hangs on the emulator, excluded in CI, no real creds
AzureServiceBus 9 Pass + FR-5 / FR-9 Deferred — verified against the live namespace azure-ci already uses. FR-5 is a platform difference (ASB dead-letters natively; there is no separate invalid-message channel), FR-9 a real defect, raised as #4318

46 of the 264 cells are Deferred, each carrying a signed-off #4240 marker.

Transport src changes (deliberately localized)

  • Paramore.Brighter.MessagingGateway.RocketMQ — guard the Baggage property; empty baggage crashed every send.
  • Paramore.Brighter.MessagingGateway.GcpPubSub — GcpPullMessageConsumer.Receive/ReceiveAsync ignored timeOut and long-polled, blocking the pump; now bound via CallSettings expiration. A null timeout returns null settings, leaving the client's own per-method default in force — Expiration.None would have overridden it with no deadline at all.
  • Paramore.Brighter.MessagingGateway.AWSSQS (V3 + V4) — sync SendWithDelay passed TimeSpan.Zero.
  • Paramore.Brighter.MessagingGateway.MQTT — Receive returned immediately on an empty buffer, so every caller needed an external sleep. The MQTT event handler now releases a semaphore per message: Receive blocks on it, ReceiveAsync awaits it and honours the caller's cancellation token. The producer takes an InstrumentationOptions constructor parameter, as every other gateway producer does, rather than always tracing at All.
  • Paramore.Brighter.MessagingGateway.RMQ.Async — RmqSubscription's isDurable default changes from false to true (both constructors). See the behaviour-change note below.

Everything else is test-harness, generator template, or ledger work.

⚠️ One user-visible behaviour change — RMQ.Async subscriptions now default to durable queues

RmqSubscription/RmqSubscription<T> in Paramore.Brighter.MessagingGateway.RMQ.Async change
bool isDurable = false to bool isDurable = true (RmqSubscription.cs:115 and :180).

Why. RabbitMQ 4.3 blocks transient non-exclusive queues by default, so the previous default
makes every affected subscription fail to declare against a 4.3+ broker — #4355.

What it means for existing users. A caller who never passed isDurable and relied on the
transient default will now get a durable queue: queue definitions survive broker restart, and
the declare will fail with a PRECONDITION_FAILED channel error against an existing queue that
was declared non-durable
, since durability is immutable after declaration. Callers who want the
old behaviour should pass isDurable: false explicitly. Nothing else about the subscription changes.

The change is deliberately asymmetric, and #4355 stays open. RMQ.Sync's RmqSubscription
still defaults to false (:106, :165). RMQ.Sync is independently non-functional against
RabbitMQ 4.3 — a measured 46 failures, all on transient_nonexcl_queues, reproduced unchanged on
a pre-branch commit
, so it is pre-existing rather than this PR's doing. Fixing it is out of scope
here; #4355 tracks the remainder. Until then, "the Brighter default" differs between the two RMQ
transports
and should not be quoted without saying which.

Terminal cleanup — done

Gated on the zero-Unknown ledger, and performed in the order the acceptance criteria require (the
test-configuration.json keys last, so each earlier step was provably a no-op):

  • the four legacy templates and their generated copies deleted — 8 templates + 72 copies, 80 files;
  • the SkipTest gate branches and the LegacyGatedTemplates closed list stripped;
  • the three retired properties removed from MessagingGatewayConfiguration;
  • the three gate keys removed from every test-configuration.json (9 files, 63 occurrences).

Regeneration after each step was a content no-op — not one generated file differs — which is the
proof that nothing was still reading those keys. The three retained capability gates
(HasSupportToPublishConfirmation, HasSupportToValidateBrokerExistence,
HasSupportToValidateInfrastructure) keep their meaning and are still tested.

One retained gate was found to be mis-declared in the same way the retired ones were, and is
corrected here rather than left. Kafka / Consumer was the only Kafka configuration setting
HasSupportToValidateInfrastructure: false, so Classic and PartitionKey ran assume_channel and
validate_channel while Consumer ran neither. Measured against a live broker, that configuration
splits cleanly: validate_channel passes both variants, assume_channel fails both,
deterministically — the flag was suppressing two working tests in order to suppress two that do not
work. A narrower gate, HasSupportToDetectMissingInfrastructureOnAssume, now skips assume_channel
alone; HasSupportToValidateInfrastructure is unchanged in meaning, so MQTT and Redis are unaffected.
The Kafka suite goes 188 -> 190 pass / 0 fail, by exactly the two tests recovered.

The underlying cause is a real behavioural difference, declared rather than papered over: EnsureTopic()
returns immediately on OnMissingChannel.Assume without an admin call, and the test compose disables
topic auto-creation, so the topic genuinely does not exist. Captured against a live broker, the classic
consumer raises ChannelFailureException wrapping ConsumeException: Subscribed topic not available: <topic>: Broker: Unknown topic or partition, while the KIP-848 consumer returns MT_NONE with no error
at all.

Fixing that error surfacing in KafkaMessageConsumer is deliberately not attempted here — it sits
outside this PR's localized src boundary — so it is raised as #4299 instead. The new flag is a
workaround for that defect, not a statement of a platform limit, and the code comment says so: when
#4299 is fixed, drop the flag from Kafka / Consumer and let the default return to true.

CI audit — no silent skips

A read-only, network-free audit (it never queries the issue tracker) enforces the deferral trail in
three ways:

  • every Skip in a canonical template or a generated Skip must match Deferred: #<n> — a bare or
    reasonless Skip fails the build. Hand-written gateway tests are out of scope by design: they are
    not generated, so a Skip on one is a maintainer's explicit choice, not something the ledger
    emitted;
  • every Deferred ledger cell must carry both an issue link and a recorded sign-off;
  • every generated test's Skip must agree with its own (configuration × behaviour) ledger cell, in
    both directions — a cell flipped without regenerating fails, whether that leaves a test wrongly
    skipped or wrongly running. The expected value comes from the generator's own emitter, so the audit
    cannot drift from what generation produces.

The audits assert their own non-vacuity (files scanned, markers found, and that both branches were
exercised), and each failure mode is covered by a synthetic-tree test, so they cannot pass by checking
nothing.

Out of scope

Fixing all brokers that don't currently implement Brighter support when there is no native support is out of scope for this work.

Every Deferred cell resolves to #4240, this PR's own tracking issue, by maintainer ruling — no
per-deferral follow-up issues are raised. The #NNNN pre-audit placeholders are fully reconciled: none
remains in the ledger or in any generated Skip marker.

🤖 Generated with Claude Code

iancooper and others added 30 commits July 18, 2026 14:40
…or, the point is not to test the details of these and foucs on the tests use of the channel and producer.
…e retirement, and rollout governance

ADR 0066 extends the generated messaging-gateway provider interfaces (DLQ +
invalid-message routing keys, invalid-channel read, in-memory and spy
scheduler-backed producers, strongly-typed RejectionMetadataKeys), retires the
HasSupportToDelayedMessages/HasSupportToDeadLetterQueue/HasSupportToRequeue
opt-in gates, and deletes the broken with_delay requeue template.

ADR 0067 sequences the rollout as fix-then-flip (Kafka reference, then the
DLQ-ADR transports, then GCP/Azure gaps, with the 0066 flip merged last) and
governs deferrals via a conformance ledger cross-checked against a mandatory
linked-issue Skip convention.

Design phase for spec 0036 (issue #4240).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Re-review of the design phase raised six findings. Amendments:

- 0066: FR-3 scheduler moved from a standalone producer factory to channel-level
  members (CreateChannelWithInMemoryScheduler / CreateChannelWithSpyScheduler ->
  SpyScheduledChannel), since the runtime seam is the consumer that backs the
  channel (ADR 0039); a standalone producer is not on the channel.Requeue path.
- 0066: specify what CreateChannelWithInMemoryScheduler actually requires
  (CommandProcessor + FireSchedulerMessage handler + external bus/producer
  registry + TimeProvider/id funcs/conflict policy) and record that cost under
  Consequences -> Negative.
- 0066: note the deliberate departure from FR-1(4)/AC-1's literal wording; fix
  stale producer-level phrasing in the RDD role bullets; name sibling ADR 0067;
  broaden the mis-declared-gate inventory (Kafka mis-declares all three).
- 0067: correct RMQ - it has no per-transport DLQ ADR (native DLX + universal
  0047/0045), so its fix may be larger; ledger rows are now per gateway
  configuration (~20) rather than per test project (9); CI audit scoped to
  in-tree artifacts, issue-state/sign-off left to the maintainer gate.
- requirements.md: amend FR-1(4), AC-1, AC-3 and the FR-3 example so the
  producer-vs-channel scheduler wording matches the design.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ic FR-2

Design review round 3 established that the scheduler seam
(IAmAChannelFactoryWithScheduler) is implemented by only six gateways - within
the generator's target set, 6 of ~20 gateway configurations. The other 14 (AWS
x4, AWS.V4 x4, GCP x4, PostgreSQL, RocketMQ) take no scheduler and delay
natively. A scheduler-delegation assertion would therefore fail by design on
conformant transports, and giving them the seam would need a public runtime API
change that C-1 forbids. It is also a mechanism assertion, which NFR-3/OOS-1
forbid.

- requirements.md: FR-2 restated as mechanism-agnostic (delayed requeue
  redelivers after the delay, regardless of native/producer/scheduler); FR-3
  withdrawn and folded into FR-2; FR-1(4) and NFR-4 withdrawn as moot; AC-2
  broadened, AC-3 withdrawn; corrected the false claim that RocketMQ is the only
  configuration declaring HasSupportToDelayedMessages true (AWS SqsStandard
  declares it too, in both V3 and V4).
- 0066: removed the scheduler-carrying provider members and spy types; added a
  "Why there is no scheduler member" section; recorded the InMemoryScheduler
  harness cost for the OOS-2 follow-up; specified the MT_NONE-on-empty read
  contract for the DLQ/invalid-channel reads so the AC-5/AC-18 negative
  assertions are writable; rewrote Alternative 4 against the seam-coverage
  evidence.
- 0067: FR-3 removed as a ledger column, with the rationale that an N/A(native)
  cell would have reintroduced the native/non-native distinction OOS-1 rejects;
  gate inventory corrected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review round 4 found two High issues, both introduced by the previous round.

1. The claim that all 14 non-scheduler configurations "delay natively" was
   generalized from the two that were verified (SQS, PostgreSQL) and is false
   for five of them. GcpPullMessageConsumer.Requeue and
   GcpPubSubStreamMessageConsumer.Requeue ignore the delay argument outright
   (the XML doc states it is "not used by Pub/Sub"; redelivery timing comes from
   the subscription RetryPolicy), and RocketMessageConsumer.Requeue is a no-op
   returning true with its ChangeInvisibleDuration call commented out pending an
   upstream RocketMQ C# client fix. Corrected in 0066 and requirements.md, and
   0067 now seeds GCP x4 and RocketMQ into the ledger as known FR-2
   non-conformances (RocketMQ flagged as a likely signed-off Deferred row, being
   blocked on a third-party dependency) rather than discovering them at the flip.

2. The FR-3 withdrawal had not reached Consequences, Risks, Alternative 3,
   References or several spots in requirements.md. Two of those were live
   instructions: 0066's Negative bullet told implementers a provider must supply
   "a scheduler-backed channel", and its 0067 reference told the ledger to track
   "the in-memory scheduler arm" - neither exists. Swept all stale FR-2/FR-3
   pairings, corrected OOS-3's transitive-proof justification (scheduler
   forwarding is no longer proven transitively), and corrected the Kafka
   coverage-gap paragraph.

Also: AC-1 reworded to match the single-CreateSubscription-with-nullable-keys
shape rather than demanding "separate members"; read-member contract extended to
cover reading a channel the subscription does not configure; MSSQL's
HasSupportToDeadLetterQueue:false added to the mis-declared inventory; and
Alternative 4 no longer conflates six gateways with six configurations.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…n scope

Rewrite requirements.md on the principle that requirements assert WHAT must be
true and how it is verified, while the ADRs carry WHY and HOW. The document had
accumulated four rounds of inline amendment scar tissue -- withdrawal markers,
"originally worded" notes, ADR rationale and a line-number evidence list -- and
was carrying a running argument instead of a specification.

Identifiers are not renumbered. FR-1(4), FR-3, FR-18, NFR-4, AC-3, AC-19 and
OOS-6 are retired as permanent gaps, preserving ~130 cross-references in ADRs
0066/0067. decision-log.md records why each was withdrawn, keeping that
deliberation out of the spec itself.

Scope: all twelve src/Paramore.Brighter.MessagingGateway.* transports are in
scope, not the nine that happen to have generator wiring. FR-20 onboards
AzureServiceBus, MQTT and RMQ.Sync (config + provider + CI infrastructure); OOS-6,
which had excluded them, is withdrawn. A missing test-configuration.json describes
what the generator covers, not what a transport owes -- the same error in kind as
gating a universal obligation behind a capability flag.

Substantive corrections from adversarial review rounds 4-6:

- FR-12/AC-12: the blanket "no template may call Requeue without a non-null
  TimeSpan" was self-contradictory -- FR-10 preserves the plain-requeue template
  and FR-15 requires Requeue(M, null). Narrowed to delayed-requeue templates.
- FR-19/AC-22: delete the requeue-count-exhaustion template. Exhaustion is
  enforced by the message pump (Message.HandledCountReached has two callers, both
  in Reactor/Proactor) or by native redrive (AWS pairs requeueCount: 3 with a
  RedrivePolicy). Channel.Requeue counts nothing. A pump test (OOS-5) or a
  native-mechanism test (NFR-3/OOS-1), so not a channel obligation.
- FR-1(6)/AC-1: remove bool setupDeadLetterQueue from CreateSubscription; a
  boolean cannot express the DLQ-only/invalid-only/neither combinations FR-1(2)
  needs. Breaking change across all 20 providers and both interface templates.
- AC-12/AC-22 also require the 38 checked-in generated copies to go; deleting a
  .liquid template does not delete its generated output.
- AC-20/AC-21 give NFR-2 and NFR-3 acceptance criteria. AC-20's exemption is per
  assertion, not per AC, so negative assertions stay unretried while the positive
  arrival half of the same criteria keeps its bounded retry loop.
- FR-13 defines the target set; C-1 widened to cover FR-20's test-side work.

Also corrects the false claim that Azure/ASB was "currently partial" (it had no
test-configuration.json at all), and ADR 0066's scheduler-seam coverage, which
counted four targeted gateways where all six are now targeted transports.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ign r5

The gates are no longer retired up front. Canonical templates become ungated
by construction, the four legacy gated templates stay suppressed until they
are deleted, and the gates and config keys retire last as a terminal cleanup.

Design review round 5 found that the old sequencing could not execute: the
flip gate required every ledger cell resolved before the gates were removed,
but while the gates are live SkipTest suppresses any template whose filename
contains requeuing, with_delay, delayed_message or dead_letter_queue -- and
Kafka, the reference transport, declares all three gates false. Its canonical
FR-2 and FR-9 tests could not generate until the very change the ledger was
meant to authorise.

The spec owner's ruling dissolved it at the root: the old tests are never
wanted, before or after the canonical set exists. A gate suppressing a legacy
template is doing useful work until that template is deleted, so retiring the
gates first would generate precisely the tests this spec replaces, against
transports not yet fixed.

Requirements:
- FR-10 rewritten as a four-part gating lifecycle; SkipTest consults the gates
  only for a closed list of four legacy template filenames, so a canonical
  template cannot be suppressed however it is named (naming cannot be relied
  on -- NFR-1 means a canonical delayed-requeue template contains both
  requeuing and with_delay).
- FR-9 now requires a canonical delayed-send template rather than ungating the
  legacy one; FR-22 + AC-25 added for canonical plain requeue, the behaviour
  old FR-10 supplied by ungating.
- FR-11 resequenced (removing a key early would ungate its legacy template);
  FR-12/FR-19 deletions folded into the legacy sweep; AC-10 became three
  ordered checkpoints; AC-9, AC-12, AC-13, AC-22 updated.
- Deletion scope is 80 generated copies across four templates, not 38 across
  two, plus 40 generated provider-interface copies broken by FR-1(6).

Requirements review round 7 (10 findings at threshold, 0 critical):
- restored FR-13's truncated definition of "targeted gateway configuration"
- FR-21 + AC-24: the conformance ledger now has a requirement
- bounded FR-20(3) to execution against a broker, not compilation
- twelve-row gateway->test-project mapping table (five pairs differ by name)
- ADR 0066: stale "~20 target" phrasings, FR-20 coverage, narration removed
- ADR 0067: C-1 widened to permit FR-20, DLQ ADRs cited by slug not number

Design review round 5 (12 findings, 9 at threshold, 0 critical): nine
remediated here, four still owed -- the ledger cannot represent a deferred
FR-20 onboarding, Scope ownership of FR-19/20/21, RejectionMetadataKeys has no
emitting template or namespace, and Reactor/Proactor parity needs a cell form
the vocabulary lacks.

Both ADR frontmatter summaries changed, so docs/adr/index.md is regenerated.
Rationale for the reversal is in the spec's decision-log.md.

No code changes; no phase approved. Requirements round 8 and design round 6
are owed on this text.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Urq9JbwARi9eqT2fhXsoFV
Runs the two owed review rounds on the gating-lifecycle reversal and
remediates every finding from each.

Requirements round 8 (8 findings, 4 at threshold, no Criticals): the
reversal had been written as if "never ungated" meant "never generated".
It does not — a gate suppresses a template only where declared false, and
most configurations declare these gates true, so the legacy templates keep
generating until deleted. Corrected in FR-9, FR-10(3), FR-22, AC-9, AC-10,
AC-22, AC-25, the terminology list, both ADRs, the README and the decision
log.

The conflation had a hidden consequence. FR-1(6) removes
setupDeadLetterQueue while the exhaustion template is still live for
sixteen configurations, and that template passes the flag positionally as
a bare `true` — so none of its 32 generated copies contains the parameter
name, and migrating "every generated caller" by searching for it misses
every broken call site. FR-1(6) now carries an interim obligation to edit
the template in the same change; AC-1 records that positional call sites
are not name-searchable.

Spec-owner ruling: FR-15 narrows to the explicit TimeSpan.Zero argument;
FR-22 owns the no-delay call in both spellings. Requeue's delay parameter
is optional and null-defaulted, so Requeue(m) and Requeue(m, null) were
one behaviour specified twice with two ledger columns.

Design round 6 (6 findings, 3 at threshold, no Criticals): all four owed
round-5 findings confirmed resolved — placeholder ledger rows per
un-onboarded transport (F3), Scope ownership of FR-20/FR-21/AC-24 (F4), a
Shared/ template giving RejectionMetadataKeys a home plus string.Empty for
unstamped fields (F5), and partial parity as a single Deferred cell (F9).

Round 6 also caught an arithmetic error introduced by the round-8
remediation: "three of the four legacy templates generate today" is four
of four (6 + 36 + 6 + 32 = 80). Corrected in both ADRs, requirements.md,
the README and the decision log, and annotated in the round-8 record.

Also fixed: the "three gate branches" off-by-one (SkipTest has four, keyed
on three gates); 0067's References entry still describing the superseded
flip-then-fix sequencing; the placeholder-row vs seeding-Unknown
disagreement; the AWS-family undercount in 0066's Context; and the FR-19
attribution for the last setupDeadLetterQueue caller.

Neither phase is approved. Both carry unreviewed remediation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Urq9JbwARi9eqT2fhXsoFV
…d arm

Remediate requirements review round 9 (9 findings, 1 Critical) and round 10
(PASS, 0 at threshold); requirements now approved.

The Critical: FR-2 required "redelivered after delay D" but AC-2 asserted only
that a later receive yields the body, so a gateway that ignores the delay and
redelivers immediately would pass. AC-2 gains a two-sided assertion — an
immediate receive must yield MT_NONE before D and the message must arrive after
it — added to AC-20's exemption list. This makes GCP x4 and RocketMQ fail as
generated, matching the ledger the ADRs seed. NFR-2's bound is quantified once
(500ms poll / 30s ceiling / 5s delay); FR-15 reworded to an assertable
first-iteration/elapsed check; FR-21 now names the five known non-conformances.

Also: provider parity tightened to both interfaces (FR-20(2)/AC-23/AC-14);
FR-19 draft-narration removed; hand-written test counts corrected to
31/19/15; PostgresSQL ledger token normalised; AC-24/AC-25 reordered; Out of
Scope reframed around Brighter-universal vs implementer-owned; ADRs 0066/0067
swept for the two-sided FR-2 and two-transports/five-configurations wording.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Urq9JbwARi9eqT2fhXsoFV
…esign

Design review round 7 (1 finding) and round 8 (PASS, 0 findings). The
round-7 finding: this session's two-sided-FR-2 edits mis-attributed
RocketMQ to FR-2's before-D (immediate-MT_NONE) arm. Verified against
source, RocketMQ's Requeue is a no-op leaving the message held by a 30s
invisibility timeout, so it passes the before-D arm; only GCP x4
(immediate redelivery) fails it. Corrected across both ADRs, requirements
(FR-2/AC-2/FR-21, AC-2 had self-contradicted), and the decision log.

Approve design: ADRs 0066 and 0067 flipped Proposed -> Accepted in both
frontmatter and body; docs/adr/index.md regenerated; .design-approved
marker added. .requirements-approved stands (factual correction only).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Urq9JbwARi9eqT2fhXsoFV
Add the 59-task unattended (ralph) implementation plan for the universal
transport conformance tests, its adversarial review record, and the tasks
approval marker.

The plan was reviewed across four rounds; the final round PASSed with zero
findings at or above threshold. The last remediation fixed:
- Brighter.sln -> Brighter.slnx in the solution-build gates
- Phase 1 "generate everywhere" now runs a structural test, not just a build
- broker startup decoupled ({ docker compose up -d || true; }) so an infra
  block still reaches the flag-and-move-on deferral gate
- Phase 3/4 multi-config tasks scoped per configuration namespace
  (AWS x4, AWS.V4 x4, GCP x4, RMQ.Async x2) so a sibling cannot fail the row
- Phase 0 exhaustion-template edit made unambiguous (pass deadLetterRoutingKey)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Urq9JbwARi9eqT2fhXsoFV
- Test: conformance-status.md artifact; RALPH-VERIFY grep gate passes
- Implementation: 23-row × 11-behaviour ledger, all cells Unknown; 3 placeholder rows; 5 known-FR-2-gap cells annotated
- Ralph task: 1/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: When_gate_flags_are_false_should_skip_only_legacy_templates
- Implementation: LegacyGatedTemplates allow-list gates the four legacy branches only; canonical templates ungated by construction
- Ralph task: 2/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
…Q bool

- Test: When_generating_provider_interface_should_expose_canonical_surface
- Implementation: interface templates gain canonical surface (routing-key params, GetMessageFromInvalidChannel, RejectionMetadataKeys, MT_NONE contract); exhaustion templates pass deadLetterRoutingKey explicitly
- Ralph task: 3/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: When_generating_gateway_should_emit_rejection_metadata_keys_once_per_config
- Implementation: new Shared/RejectionMetadataKeys.cs.liquid + once-per-config emit in MessagingGatewayGenerator; csproj copies Shared templates; disable assembly test parallelization to prevent shared-template-dir race
- Ralph task: 4/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.Kafka.Tests succeeds
- Implementation: both Kafka providers implement routing-key CreateSubscription, GetMessageFromInvalidChannel, RejectionMetadataKeys (PascalCase keys); regenerated Generated tree (interface copies + Shared record); no setupDeadLetterQueue remains
- Ralph task: 5/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.AWS.Tests succeeds
- Implementation: four AWS providers (Sns/Sqs × Standard/Fifo) implement routing-key CreateSubscription, GetMessageFromInvalidChannel(+Async), RejectionMetadataKeys (SQS camelCase keys); regenerated Generated tree; no setupDeadLetterQueue remains
- Ralph task: 6/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.AWS.V4.Tests succeeds
- Implementation: four AWS.V4 providers implement routing-key CreateSubscription, GetMessageFromInvalidChannel(+Async), RejectionMetadataKeys (SQS camelCase keys); regenerated Generated tree; no setupDeadLetterQueue remains
- Ralph task: 7/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.Gcp.Tests succeeds
- Implementation: four GCP providers (Pull/PullOrdering/Stream/StreamOrdering) implement routing-key CreateSubscription, GetMessageFromInvalidChannel(+Async), RejectionMetadataKeys (string.Empty where unstamped); regenerated Generated tree; FR-2 gap deferred to Phase 4
- Ralph task: 8/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.MSSQL.Tests succeeds
- Implementation: MsSqlMessageGatewayProvider implements routing-key CreateSubscription (explicit deadLetterRoutingKey), GetMessageFromInvalidChannel(+Async), RejectionMetadataKeys; regenerated Generated tree; no setupDeadLetterQueue remains
- Ralph task: 9/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
…tgresSQL

- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.PostgresSQL.Tests succeeds
- Implementation: PostgresMessageGatewayProvider implements routing-key CreateSubscription, GetMessageFromInvalidChannel(+Async), RejectionMetadataKeys; regenerated Generated tree; no setupDeadLetterQueue remains
- Ralph task: 10/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.Redis.Tests succeeds
- Implementation: RedisMessageGatewayProvider implements routing-key CreateSubscription, GetMessageFromInvalidChannel(+Async), RejectionMetadataKeys (Redis camelCase keys); regenerated Generated tree; no setupDeadLetterQueue remains
- Ralph task: 11/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
…Async

- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.RMQ.Async.Tests succeeds
- Implementation: RmqClassic/Quorum providers implement routing-key CreateSubscription, GetMessageFromInvalidChannel(+Async), RejectionMetadataKeys (string.Empty; RMQ uses native DLX); regenerated Generated tree; no setupDeadLetterQueue remains
- Ralph task: 12/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: compilation gate (AC-1); dotnet build tests/Paramore.Brighter.RocketMQ.Tests succeeds
- Implementation: RocketMqMessageGatewayProvider implements routing-key CreateSubscription, GetMessageFromInvalidChannel(+Async), RejectionMetadataKeys; regenerated Generated tree. Last of the 20 provider migrations — no actual setupDeadLetterQueue parameter remains in the repo.
- Note: the task RALPH-VERIFY's repo-wide 'grep -rn setupDeadLetterQueue tests tools' has a benign false-positive — the only two matches are Assert.DoesNotContain("setupDeadLetterQueue", ...) absence-assertions in the FR-1 interface meta-test (task 3). RocketMQ builds clean and no real parameter/usage remains. Flagged for owner (candidate Phase 6 grep scoping).
- Ralph task: 13/59

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
… Phase 6 grep

- ADR 0066: add 'Implementation notes (learned during Phase 0 rollout)' under
  Implementation Approach (status unchanged/Accepted) — regen needs --framework net10.0
  (generator multi-targets); generator-test project must disable xUnit parallelization
  (shared Templates output dir race); repo-wide absence greps collide with meta-tests
  that name the token.
- ralph-tasks.md: add --framework net10.0 to all 35 regenerate RALPH-VERIFY commands and
  the two Execution Notes describing them (the bare dotnet run aborts on a multi-targeted
  generator).
- ralph-tasks.md Phase 6: harden the #NNNN reconciliation grep to exclude **/ConformanceAudit/**
  so audit negative-fixtures naming #NNNN don't false-positive (same class as the Phase 0
  setupDeadLetterQueue grep vs the FR-1 interface meta-test's Assert.DoesNotContain).

Co-Authored-By: Claude Opus <noreply@anthropic.com>
…l templates

- Test: When_ledger_marks_a_cell_should_emit_skip_only_when_not_proven
- Implementation: load+parse conformance ledger (repo-root resolution via walk-up),
  canonical-template->FR-column map, per-cell Skip value into render context; empty
  Skip suppressed via Liquid empty-string equality; canonical templates only
- Ralph task: 14/59 (Phase 1 task 1)

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
…livery and ledger-driven Skip

- Test: When_generating_plain_requeue_should_emit_bounded_redelivery_both_variants
- Implementation: Reactor + Proactor canonical templates emitting Requeue/RequeueAsync
  no-delay, asserting true and bounded 500ms-poll/30s-ceiling redelivery loop (AC-20);
  conditional ledger-driven Skip, no hard-coded marker. Also fixed task-1 test isolation
  to restore (not delete) the canonical template in Dispose
- Ralph task: 15/59 (Phase 1 task 2)

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
…r-D arms)

- Test: When_generating_requeue_with_delay_should_emit_before_and_after_arms_both_variants
- Implementation: Reactor + Proactor canonical templates passing a positive 5s TimeSpan to
  Requeue/RequeueAsync; before-D arm single immediate receive asserts MT_NONE (AC-20 exemption),
  after-D arm asserts arrival inside the bounded 500ms/30s retry loop; no mechanism assertion
  (AC-21); conditional ledger-driven Skip, no hard-coded marker
- Ralph task: 16/59 (Phase 1 task 3)

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: When_generating_zero_delay_requeue_should_emit_first_iteration_receipt_both_variants
- Implementation: Reactor + Proactor canonical templates calling Requeue(M, TimeSpan.Zero)
  explicitly; asserts true, first-iteration receipt inside the bounded retry loop, and
  elapsed-under-5s (proves zero is neither special-cased nor a positive delay); conditional
  ledger-driven Skip, no hard-coded marker
- Ralph task: 17/59 (Phase 1 task 4)

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
- Test: When_generating_delivery_error_reject_should_emit_dlq_routing_both_variants
- Implementation: Reactor + Proactor canonical templates proving Reject(M, DeliveryError)
  on a channel with a deadLetterRoutingKey routes M to the DLQ; asserts original-topic
  (== data topic) and rejection-reason via bounded GetMessageFromDeadLetterQueue; conditional
  ledger-driven Skip, no hard-coded marker
- Ralph task: 18/59 (Phase 1 task 5)

Co-Authored-By: Claude Opus <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet <noreply@anthropic.com>
iancooper and others added 21 commits September 17, 2026 13:57
…orbids

The two-message templates built both messages from one DefaultMessageBuilder,
whose id and body are field initialisers - so the "two" messages shared an id
and a body, and were the same message. The terminal assertion was only
Assert.NotEqual(MT_NONE, ...), which a redelivered first message satisfies.

That is the one outcome the behaviour exists to forbid: a rejection with no
dead-letter queue and no invalid-message channel must REMOVE the message, not
redeliver it. The test could not tell removal from redelivery.

Both messages now carry their own id and body, and the test asserts the
identity of what arrives. Order matters: the assertion alone would have changed
nothing, because with equal ids it passes either way.

Proved in both directions - shared ids fail the build-site count, and removing
the identity assertion while keeping distinct ids fails too, so neither half is
vacuous.

Variables renamed from message1/message2/received1/received2 to
rejectedMessage/followingMessage/received/receivedFollowing.

Generator suite 269/269, solution builds with 0 errors. 48 generated files
regenerated. FR-16 carries the same defect and follows as its own change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…r is

The previous pass kept the FR-n tokens that name conformance-ledger columns,
on the grounds that the code really does read those columns. That was right
about the tokens and wrong about the prose: "its FR-8 cell claims Pass" tells
a reader nothing unless they already know what an FR-8 cell is and where to
find one.

Each of the 43 now names the behaviour and says where the column lives:

  - "its FR-8 cell claims Pass"
    -> "the conformance ledger's rejection-metadata column (FR-8) claims Pass"
  - "FR-16 is the \"Nack redelivers\" column"
    -> "the Nack-redelivers behaviour is column FR-16 of the conformance ledger"
  - "the generator keys on an FR-2 + FR-4 column pair"
    -> "the generator keys on the conformance ledger's FR-2 + FR-4 column pair"

The reconciliation table needed only its header changed - saying once that the
right-hand column is "the behaviour column it is judged against in the
conformance ledger" explains all eleven rows.

Behaviour names are taken from CanonicalBehaviours.FR_COLUMN_BEHAVIOURS, so the
prose and the generator's own labels cannot drift apart.

Structural only: every changed line is a comment. Generator suite 269/269,
solution builds with 0 errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The same defect as the no-channels rejection test, and the second half of the
blocker. The two-message fact built both messages from one DefaultMessageBuilder
whose id and body are field initialisers, so the "two" messages were one, and:

  - "the nacked message comes back before the one queued behind it" passed
    whichever message actually arrived;
  - the following message's arrival was only Assert.NotEqual(MT_NONE, ...),
    which the nacked message arriving a second time satisfies.

Every build site now states its own id and body - the single-message fact
included, since the builder is a shared field and its safety there was
incidental rather than designed - and the test asserts the identity of the
message that follows.

Proved in both directions: shared ids fail the build-site count, and removing
the identity assertion while keeping distinct ids fails too.

Two pre-existing facts asserted Assert.Contains("message2") as a proxy for "the
two-message variant exists". That coupled a test to an incidental variable
name; they now assert the fact's own name. The last M1/M2 spellings in the
generator suite go with them, including a test method that carried m2 in its
name.

Variables renamed to nackedMessage/followingMessage/receivedForNack/
receivedFollowing.

Generator suite 271/271, solution builds with 0 errors, 48 generated files
regenerated. No canonical template now takes two messages from one builder's
defaults.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
AC-7 and AC-17 mandated a delivery order — "the next receive yields M2",
"the redelivered M is received and then M2". Ordering is a property of the
transport, not a Brighter behaviour, and the targeted transports do not
share one: SNS/SQS Standard and GCP Pub/Sub without an ordering key
guarantee nothing. The ACs therefore made a conforming gateway untestable
on an unordered transport, and FR-7's Example re-imported the order that
the FR prose had deliberately left out.

Amends AC-7, AC-17 and FR-7's Example to identify messages by id rather
than by arrival position, and adds a clause to NFR-2 permitting a bounded
loop to drain to a named target.

Adds NFR-4 (Ordering neutrality): a multi-message arm identifies receipts
by id and asserts over the set observed; a redelivery is proven by seeing
an id a second time after the release, never by its position; arms
tolerate repeat receipts because transports are at-least-once.

NFR-4 also records that a guarantee the transport does offer is void
across a delay: the scheduler seam honours a delay by re-publishing, and
a re-publish is a new enqueue, so it lands at the tail of the queue,
partition or message group even where ordering is guaranteed.

The per-configuration ordering table is stated as owed rather than
guessed — only five rows are verified against provider code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The FR-16 nack-redelivery and FR-7 no-channels-rejection arms paired a
named sent message with whichever message the transport happened to
deliver first. The assumption was latent for as long as both messages
were byte-identical; febd0d7 and 2e650c3 gave each its own id and
body to fix a real blindness, and that exposed it. aws-ci then went red
on SNS/SQS Standard only, non-deterministically, with the identity
assertion receiving the test's other message — while every FIFO variant
passed, because FifoMessageBuilder stamps one partition key per builder
instance and so puts both messages in one ordered, head-of-line-blocked
group.

Root cause: the arms asserted a delivery order that SQS/SNS Standard does
not provide and that FR-7 and FR-16 never required — only their
illustrative examples did.

Both arms now resolve identity by id against the messages they sent and
assert over the set of ids observed within the NFR-2 bound:

- FR-16 nacks whichever message arrives first, then drains one bounded
  loop until both ids are seen, acknowledging each receipt. Seeing the
  nacked id a second time is a redelivery by definition, because that
  message was already received once and released — so an implementation
  that never redelivers still fails, and the guarantee 2e650c3 restored
  is kept. Acknowledging as it drains is what lets a transport that
  blocks a message group make progress.
- FR-7 rejects whichever message arrives first, still forbids the
  rejected message coming back, and asserts full identity of the other.

The single-message nack arm is untouched: with one message in flight no
ordering question arises, so its positional assertion remains correct.

Templates are the change; the 96 regenerated files are its output. The
four golden assertions that pinned the old positional expressions move
with them, and a new generator test pins the rule broker-free.

Generator suite 275/275, dotnet build Brighter.slnx 0 errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
aws-ci is green on f7ad18d with every canonical arm RUN — 0 skipped,
0 failed, across both SDK majors and both variants, twice.

This is the first evidence that FR-7 and FR-16 hold on an unordered
transport: before the fix the two messages were first indistinguishable
and then distinguishable but order-dependent, so no test could tell. The
four AWS{,.V4} / Sns|SqsStandard cells are now earned rather than
inherited and need no ledger change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The contributor docs had no route into the test generator, and what they did
say about it was wrong in three ways:

- "Generated tests can be customized after generation" contradicted the
  never-edit-a-generated-file rule and would be destroyed by the next
  ./generate-test.sh. Replaced with the three real seams: the Liquid template,
  the hand-written provider class, and test-configuration.json.
- The single-project invocation was the one the reference doc explicitly
  labels "Wrong" - it generates into the repo root, because the generator uses
  CWD as its output root. Replaced with build-then-cd, which also covers the
  stale-template trap.
- The ADR 0035 link 404'd on a misspelled filename.

TESTING_GUIDE.md has a fully generic name but is entirely about RabbitMQ mTLS,
so a newcomer looking for a testing guide finds it first. Renamed to
RABBITMQ_MTLS_TESTING_GUIDE.md, scoped in its opening lines, and its
"repository root" hard-coded to one contributor's home directory made generic.
Nothing in the repo linked to it.

Added a README to the generator itself, which had none: what it generates, the
never-edit rule, how to run it, the twelve canonical behaviours with their
ledger columns, how a ledger cell decides whether a test runs or is skipped -
which is how a contributor finds work to pick up - and links to the full
reference and ADRs 0035/0066/0067.

Docs only; no template, generated file or production code is touched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The knowledge a contributor needs about the conformance suite exists, but it
is organised as reference and buried where nobody looks: 482 lines in an
agent-instructions folder, and an 842-line spec whose capability matrix starts
at line 815 and keys its columns FR-2..FR-23 rather than by behaviour. Three
questions had no answer anywhere: what does the suite prove, how do I add a
transport, and how do I add a behaviour.

Three guides in docs/guides/, modelled on the box-provisioning pair:

- getting-started: the twelve behaviours by name, Reactor/Proactor parity and
  why both must pass, and - the part most often misread - what the suite
  deliberately does NOT prove: no mechanism assertions, no ordering
  assumptions, no pump, bounded rather than exact timing. Then how to read the
  matrix, how to run one transport locally, and how a "Deferred -> #NNNN" cell
  is a work item rather than a wart. That last is the "where do I pick up
  work" answer and was previously invisible.

- new-transport: NATS end to end. test-configuration.json, the provider
  surface member by member, the two rules that get broken most often (helpers
  must not retry internally; rejection metadata is all-or-nothing), the ledger
  row starting as Unknown, generation, docker-compose plus CI job, then
  driving cells. Ends with the eight audits and what each one means when it
  fails you.

- new-behaviour: the three tests for whether something is canonical at all,
  then every file a thirteenth behaviour touches, including the two hard-coded
  name lists that shadow CanonicalBehaviours and under-check silently rather
  than failing. Uses bugfix 0021 as the worked example, because a real loop
  teaches what an invented one cannot - not least that an *Example* in a spec
  gets implemented, and that local green does not mean conformant.

Linked from CONTRIBUTING.md and the generator README so they are reachable
from where a contributor actually starts.

Every number, link and quoted string was checked against the tree: all
relative links resolve, the quoted Skip string is byte-identical to a real
generated one, the provider surface is read from the generated interface, and
the 24 configurations / 24 ledger rows were counted rather than recalled.

Docs only. Generator suite 275/275 on net9.0 and net10.0, unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The Generated Tests section described the generator only as a way to keep
outbox/inbox implementations consistent. A contributor reading it would not
learn that transport conformance exists at all, nor that there is a matrix
saying what each transport is known to do, nor that its Deferred cells are a
work queue with filed issues behind them.

Rewrote the opening to name both families, say what each proves and where,
point at the conformance matrix and explain how the generator reads it, and
lead with the three guides. Chosen over adding the same pointer to the root
README: CONTRIBUTING is where a contributor is already sent, and one front
door is better than two that can disagree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FR-15 and AC-16 said "elapsed time from the Requeue call to receipt". Read
literally that starts the clock before the call, so the budget is charged for
the round trip that issues the instruction to the broker -
ChangeMessageVisibility, a nack, a list push - as well as for the redelivery
it is meant to measure.

That is the wrong thing to measure. The call's duration is the cost of asking;
the obligation is about what the gateway does with the zero once asked. On a
remote broker the ask alone can consume the whole 5 s budget while the gateway
behaves perfectly - aws-ci failed exactly this way on 194522a, missing by
976 ms on a commit whose only change was a markdown file.

Both now say "from the return of the Requeue call", and say why the call's own
duration is excluded. The 5 s figure is unchanged: this narrows what is being
timed, it does not relax the bound.

The template change that implements this follows separately.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The zero-delay requeue templates started the stopwatch BEFORE calling
Requeue, so the 5 s budget paid for the round trip that issues the instruction
to the broker - ChangeMessageVisibility on SQS, a nack, a list push - as well
as for the redelivery it exists to measure.

That is the wrong thing to measure. FR-15 asks whether the gateway treats an
explicit TimeSpan.Zero as a delay; the cost of asking it to is the harness's
network latency, not the gateway's answer. On a remote broker the ask alone
can consume the whole budget while the gateway behaves perfectly. aws-ci
failed exactly this way on 194522a - SqsStandard.Proactor missing by 976 ms
on a commit whose only change was a markdown file, with the same test passing
three other times in the same job.

The stopwatch now starts after Requeue/RequeueAsync returns. The 5 s bound is
unchanged and the 30 s loop ceiling still reads the same stopwatch, so this
narrows what is timed rather than relaxing the bound. Both assertion messages
now say "returning" so a future failure reports what was actually measured.

Pinned by two new broker-free facts in the zero-delay generator test, which
assert the stopwatch appears after the requeue call in the generated file.
Proved RED first: both failed on the ordering assertion itself - the two
index guards passed, so the strings resolved - reporting the stopwatch 54
characters ahead of the requeue.

Implements the wording landed in 3837651 (FR-15 / AC-16 now measure from
the return of the Requeue call).

Regeneration touched exactly 48 files - 24 configurations x 2 variants - and
nothing else. Generator suite 277/277 on net9.0 and net10.0; dotnet build
Brighter.slnx 0 errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
FifoMetadataProducer justified its MessageDeduplicationId stamp with a
premise that stopped being true: that the canonical suite reuses one
message builder, so its two "distinct" messages share an id and body and
FIFO would collapse them. The canonical templates have stamped a distinct
id and body per build since the two-message arms were fixed.

The stamp is still required, for a reason the comments did not give.
Every conformance FIFO queue and topic is created with
contentBasedDeduplication: false, and the gateway sets
MessageDeduplicationId only when the message bag carries one
(SqsMessageSender.SetFifoQueueProperties,
SnsMessagePublisher.ConfigureFifoSettings). Without the stamp, AWS
rejects every FIFO send in the suite — not just the two-message arms.

Both comments now say that. Comment-only: a -U0 diff filtered to
non-comment lines is empty across both files, and both projects build
with 0 errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
NFR-4 carried an explicit debt: a per-configuration statement of which
transports guarantee ordering for plain delivery, with twelve of the
twenty-four configurations recorded as "not yet verified against their
provider code" and forbidden from being called ordered until they were.

All twelve are now read off the code, and the paragraph is replaced by
the full table. Kafka Classic and Consumer are totally ordered because
both publications are single-partition and the publication defaults
(idempotence on, one in-flight request) stop a producer retry reordering
them. RabbitMQ is ordered per queue, Redis by its FIFO list, Postgres
and MSSQL by an insertion-ordered dequeue. Azure Service Bus, RocketMQ
and MQTT guarantee nothing as configured, and the table says why in each
case rather than leaving it to the transport's reputation.

Two findings are worth more than a table cell, so they are called out
under it. Kafka PartitionKey publishes to a two-partition topic and is
ordered only because FifoMessageBuilder gives one random key per builder
instance - the same harness-induced ordering that made the AWS FIFO
cells look earned before bug 0021. And RabbitMQ voids its own guarantee
on requeue at any delay, including TimeSpan.Zero, because Requeue
re-publishes before acknowledging the original; the delay clause already
above reaches that case, but by the requeue mechanism rather than the
scheduler seam.

Docs only. Generator suite 277/277 on net9.0 and net10.0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The gate that requires every wired configuration to hold the full
canonical suite kept its own copy of the canonical set, and the copy had
drifted: eleven entries against the generator's twelve. The behaviour it
omitted is the requeue-too-many-times one (FR-23 in the conformance
ledger), which the generator emits into all twenty-four configurations
and which the gate therefore checked in none of them. Both presence
facts and the ledger Skip fact iterated the stale copy, so deleting that
file from every wired project would not have turned the suite red.

The two local copies are gone; all three facts now iterate
CanonicalBehaviours.TEMPLATE_FR_COLUMNS, which is what the generator
itself reads, so "the full canonical suite" cannot mean two different
things again.

A gate scanning a complete tree cannot demonstrate what it looks for, so
the hole is pinned by a canary that hands the presence check a directory
built from the generator's map less one file. Its three facts fail and
pass in the directions that matter: before this change the FR-23 fact
reported an empty missing-list while the FR-22 control reported the
absence correctly, which is the blindness itself; the third fact keeps a
complete directory from being reported incomplete, so neither of the
others can pass by accident. Proved against the real tree too - removing
one generated FR-23 file now fails the gate by name, where it was
previously silent.

The presence loop, duplicated between the two variant facts, is now one
method so the canary can reach it.

Note the third copy of the eleven-name list, in the Kafka reference
surface reconciliation, is deliberately left alone: eleven is correct
there, because FR-23 was never part of that surface.

Generator suite 280/280 on net9.0 and net10.0, up from 277 by the three
canary facts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…or, not real Pub/Sub

Moves FR-7, FR-9, FR-15, FR-16 and FR-22 to Deferred -> #4240 for both
GCP / Pull and GCP / PullOrdering, and regenerates the twenty Skip attributes
that follow. Both rows now claim nothing, like GCP / Stream already did.

The cells were set by the Phase 4 emulator run (decision-log, 2026-07-30),
which recorded at the time that no real-GCP credentials were available. The
emulator redelivers promptly. Real Pub/Sub does not, and gcp-ci has been
saying so on every push:

  FR-15  failed 4 of 4 runs, both rows, both variants, worst 00:00:27.27
         against a 5 s bound
  FR-16  failed 3 of 4 runs
  FR-22  failed 3 of 4 runs
  FR-7   failed 1 of 4 runs
  FR-9   failed 1 of 4 runs

Every one of those assertions sits after a bounded poll loop, so none is the
harness failing to look. Where FR-16/FR-22/FR-9 pass they pass in 28 s, 24 s
and 22 s against 30 s bounds — 2 to 8 s from red, which is why the failing
set moved run to run and read as flake.

One root cause: on Pub/Sub Pull, redelivery waits for the ack deadline to
expire. ModifyAckDeadline(ackId, 0) succeeds — the test asserts the requeue
returned true before the stopwatch starts — and is ignored as an instruction
to redeliver now. It is not our own backoff: RequeueDelay defaults to
TimeSpan.Zero, so RetryPolicy is left null. It tracks the deadline instead:
the subscription sets ackDeadlineSeconds: 10 and every measured elapsed is
above 10, never below. GcpPullMessageGatewayProvider.cs:166 already said as
much for Nack.

10 s is Pub/Sub's minimum ack deadline, so a 5 s bound is unsatisfiable on
Pull whatever the harness does. FR-15's bound was not loosened: it is
normative in requirements.md FR-15 / AC-16.

FR-9 gives up a Fixed (#4240). The delayed-send fix is real and stays in the
code; the downgrade says only that real Pub/Sub Pull does not deliver it
reliably enough to claim the cell.

ADR 0067 preconditions: evidence is the actual CI assertion text over four
runs; the in-scope remedies are already in place (documented nack, bounded
polling, deadline at the service floor); the residual blocker — Cloud Pub/Sub
redelivery latency — is external.

Generator suite 280/280 on net10.0, Skip-to-ledger cross-check audit included.
Gcp.Tests builds with 0 errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…eate it

azure-ci's only two failures are the generated concurrent-post fact, Reactor
and Proactor:

  When_multiple_threads_try_to_post_a_message_at_the_same_time_should_not_throw_exception
  ServiceBusException SubCode=40901 "Another conflicting operation is in
  progress" / SubCode=40900 "not allowed in the resource's current state",
  HTTP 409, on topics named gen-asb-topic-...

Those names are minted by this harness. The test sends from four threads at
once against a topic that does not exist yet, and the gateway creates a
missing topic lazily on first send (AzureServiceBusTopicMessageProducer
.EnsureChannelExistsAsync), so four creates for the same name reach Azure
together and the losers get a 409.

The race is in the harness's arrangement, not in Brighter. The fact is about
concurrent sends; provisioning concurrently is incidental to it. So the
provider now creates the topic once before handing the producer out, which is
the producer-side twin of the EnsureSubscriptionExistsAsync it already does
for channels. Pre-creating with CreateTopicAsync(topicName) matches the
gateway's own call exactly, so no topic settings drift.

Removing the race rather than waiting out the throttle also means this does
not depend on the namespace tier, where management-operation limits live.

Guarded on MakeChannels == Create, so the two facts that assert on missing
infrastructure still find it missing — both use OnMissingChannel.Assume for
the publication and are unaffected.

Verification is azure-ci: there is no local ASB namespace, so this cannot be
run here. Solution builds 0 errors; generate-test.sh changes no file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
rocketmq-ci is commented out (ci.yml:707, "Rafael Andrade is working on how to
run RocketMQ on GHA") and the build job's project list is Test.Generator /
Core / Extensions / Transforms.Adaptors, so nothing in CI executed any part of
the RocketMQ gateway - including this branch's largest src change, the builder
block extracted out of RocketMqMessageProducer into RocketMqMessagePublisher.
RocketMqEmptyHeaderPropertyTests covers that publisher seam and needs no
broker, so there is no reason for it to sit unrun.

It could not be selected, though: all 56 tests in the project carry
Trait("Category", "RocketMQ"), the broker-dependent 47 included, so no category
filter could separate them. The class now carries a second Category value,
RocketMQBrokerFree, and the new step filters on that.

Measured here, both frameworks:

  --filter "Category=RocketMQBrokerFree"  ->  3 passed, 0 skipped, ~25 ms
  the project as a whole                 ->  56 tests
  --filter "Category=RocketMQBrokerFreeXX" -> "No test matches"

So the filter selects exactly the broker-free class, leaves the 47 that need a
broker alone, and is not matching by accident.

This covers the publisher seam, not the whole gateway. Full RocketMQ coverage
still waits on rocketmq-ci, and #4353 records that Requeue is a no-op there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The step added in 69a05e1 failed the build job:

  Could not find a test logger with AssemblyQualifiedName, URI or
  FriendlyName 'GitHubActions'.

--logger GitHubActions was copied from the Core Tests step, but that logger
comes from a per-project PackageReference which RocketMQ.Tests never had -
reasonably, since nothing in CI had ever run it. Every other test project the
workflow invokes carries the reference, and the version is already centrally
pinned in Directory.Packages.props, so this just brings the project in line.

My local check ran the filter under the default Debug configuration with no
logger, so it exercised the trait but not the step. Re-run as the workflow
actually spells it - -c Release --logger GitHubActions - it is now 3 passed on
net9.0 and net10.0 and exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…a scope extended

The umbrella ruling's scope table enumerated GCP / Pull's deferred cells as
FR-2/4/5/6/8/17, which 7bf4444 extended to all twelve. The ruling itself is
unchanged - every Deferred cell still tracks to #4240 - but the enumeration
was stale, so it now points at the reversal entry.

Also records that the root cause already has its own issue, #4321, and that
deferring the cells turns gcp-ci green, which removes the CI signal that issue
was relying on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
7bf4444's entry said measured elapsed was 'always above 10, never below'.
That is false. Across the four gcp-ci runs there are eleven measurements -
7.12, 10.21, 12.69, 13.65, 14.22, 15.21, 18.66, 18.77, 21.88, 23.73, 27.27 s
- and the first is below the 10 s ack deadline.

It is not a counter-example to the root cause, and the entry now says why:
the test's stopwatch starts when Requeue returns, while the deadline clock
started earlier, at receive, so elapsed-since-requeue can fall short of a
full deadline period. What holds without qualification is the part that
matters - every measurement is far past the 5 s bound and none is prompt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@iancooper

Copy link
Copy Markdown
Member Author

Reviewer guide: the product-code surface is 10 files, +532/−154

This PR is 963 files and ~102k lines, but almost all of it is generated tests, templates and the conformance ledger. The code that ships to users is this, and it is the whole of it:

git diff origin/master...HEAD -- src/
file ± what to look at
MessagingGateway.MQTT/MQTTMessageConsumer.cs +138 −25 the largest behavioural change. Receive used to return immediately on an empty buffer, so every caller needed an external sleep. A SemaphoreSlim is now released per arriving message; Receive blocks on it and ReceiveAsync awaits it, honouring the caller's cancellation token. Worth the closest read of anything here — it is a blocking/async change in a pump path
MessagingGateway.RocketMQ/RocketMqMessagePublisher.cs +158 −0 new file, but not new logic — it is the message-building block lifted out of the producer below
MessagingGateway.RocketMQ/RocketMqMessageProducer.cs +3 −102 the other half of that extraction. The 102 removed lines become a 3-line call. ⚠️ See the coverage caveat below
MessagingGateway.MQTT/MqttMessageCreator.cs +81 −0 new; mirrors KafkaMessageCreator / RmqMessageCreator
MessagingGateway.GcpPubSub/GcpPullMessageConsumer.cs +58 −6 Receive/ReceiveAsync only — the requeue/reject paths are untouched. Binds the pull to the caller's timeOut via CallSettings expiration and maps DeadlineExceeded to an empty message. A null timeout deliberately yields null settings so the client's own per-method default stays in force
MessagingGateway.AWSSQS/SnsMessageProducer.cs +37 −7 sync SendWithDelay passed TimeSpan.Zero
MessagingGateway.AWSSQS.V4/SnsMessageProducer.cs +37 −7 ⭐ the V3 and V4 hunks are byte-identical — verified, so reviewing one reviews both
MessagingGateway.MQTT/MQTTMessageProducer.cs +10 −1 takes an InstrumentationOptions ctor parameter instead of always tracing at All
MessagingGateway.RMQ.Async/RmqSubscription.cs +8 −4 ⚠️ the only user-visible behaviour change in the PR — isDurable defaults false → true. See the callout in the PR description
MessagingGateway.MQTT/MqttMessageConsumerFactory.cs +2 −2 wiring for the above

Three things I would not want a reviewer to take on trust

1. RocketMQ's 263 lines are barely executed. rocketmq-ci is commented out in ci.yml ("TODO: Rafael Andrade is working on how to run RocketMQ on GHA"). This PR adds a build-job step for the broker-free tests, but that covers 3 of the project's 56 tests and exercises the publisher seam only. The extraction was reviewed as behaviour-preserving by diffing the moved block — textually identical bar mqPublication→publication and connection.TimerProvider→the parameter — but CI is not proving it. The ledger's 9 Fixed RocketMQ cells rest on a one-off local run against a live RocketMQ 5.5.0 broker (21 pass / 2 skip per variant), not on anything repeatable.

2. GCP now claims nothing. All 24 GCP ledger cells are Deferred -> #4240. gcp-ci is green at 86 passed / 52 skipped — green partly by skipping. 28 gateway tests do still execute (send, receive, activity-context propagation, multi-thread post), which is what covers the Receive change above; but GCP conformance for requeue/nack/reject is now unproven, and the root cause has its own issue, #4321. Real Pub/Sub does not redeliver promptly on ModifyAckDeadline(ackId, 0); the earlier green cells came from the emulator.

3. The bot review did not reach everything. It read the FR-7, FR-16 and FR-23 templates in full and skimmed the rest of the 29; it spot-checked 3 of 24 provider files. Absence of findings in the other templates and providers is not evidence about them.

Known-flaky jobs, both pre-existing

  • #4376 — postgres-ci, hand-written gateway tests using a fixed sleep plus one read. master is 10/10 green; the branch adds load and tips them over. Neither failing test is touched by this PR.
  • #4330 — kafka-ci, Subscribed topic not available. Reproduces on master.

A red run on either is not a signal about this PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

3 - Done feature request .NET Pull requests that update .net code V10.X

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant