Repository navigation
Conversation
skycap.wandb_artifact (default false) uploads each step's {id}.json.zst
documents, without sidecars, as a version of the skycap-records-<run id>
artifact, aliased step-N and latest. Retried attempts are included. An
upload failure is logged and does not fail the step.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nment, timestamped run names MODAL_TOKEN_ID and MODAL_TOKEN_SECRET in the environment skip MODAL_KEY_FILE. EXPERIMENT can be set, and the default name carries a UTC timestamp. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… arguments The servers always write records to record_dir; wandb_artifact alone turns the upload on. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kouroshHakha
left a comment
There was a problem hiding this comment.
Using a W&B artifact for this is the right tool, and it isn't slow: the heavy upload happens in wandb's background service, the call runs off the event loop, and errors are caught, so a W&B outage can't fail rollouts. The main issue is that it uploads once per generate() call, and that isn't once per training step. Three things to fix before merging:
- Eval takes the
step-Nandlatestaliases. Eval calls the samegenerate()with the sameglobal_step, once per eval batch, right after training step N and at step 0. An alias belongs to one version at a time, so the eval upload movesstep-Noff the training records.training_phaseis inbatch_metadataand isn't checked. - Several versions per step. Dynamic sampling calls
generate()repeatedly within one step, and each call movesstep-Nto itself, so the earlier versions are reachable only asvK. - Silent partial uploads across nodes. Documents that aren't on the trainer's disk are skipped quietly. With
placement_strategy="SPREAD"and the node-local defaultrecord_dir, that's the likely case on multi-node, andnum_trajectoriesstill counts every id.
What I'd do:
- Upload once per training step, train phase only. Accumulate ids per
global_step, upload when the step changes and at shutdown, and send eval to its ownskycap-records-eval-<run>artifact or leave it off by default. - Fetch documents from the server that owns them (
GET /trajectories/{id}already reads from disk) rather than from the trainer's filesystem. The pool knows which server holds each id, so the "one node or a sharedrecord_dir" constraint goes away. - Don't await it on the step. Run it as a background task, optionally with sampling. Hashing 256 long documents costs roughly 10-100 ms per step at 32x8, which isn't big, but the step doesn't need to wait for it.
Smaller notes are inline. Size looks fine: roughly 10-80 MB per step at 32x8 with 32k-token trajectories, and memory stays flat since it streams files. A wandb.Table would have been much worse.
Eval batches call generate with the training step's global_step, so a step-N alias moved to the eval records. Uploads are now aliased train-step-N or eval-step-N, and the metadata records training_phase. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ir servers SkycapUploads uploads the train records once per step and the eval records once per eval pass, on a background thread. Each document comes from the server that wrote it. A step.json marks each attempt trained or dropped. Each phase gets its own artifact, aliased step-N and latest. The skycap.wandb group replaces skycap.wandb_artifact. CallbackInput carries the step's trajectory_ids at on_step_end. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Without MODAL_TOKEN_ID or MODAL_KEY_FILE, the script runs when ~/.modal.toml exists, and the Modal SDK reads it. That covers one node only, because SkyRL forwards Modal credentials to Ray workers from the environment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kouroshHakha
left a comment
There was a problem hiding this comment.
leaving some comments.
| record_dir: Optional[str] = None | ||
| """Where ended trajectories are written. Defaults to ``{trainer.export_path}/skycap``; each server writes | ||
| on its own node, so point it at a shared filesystem to have one directory for the run.""" | ||
| wandb: SkycapWandbConfig = field(default_factory=SkycapWandbConfig) |
There was a problem hiding this comment.
nit style: comments are not added
|
A thought on where this should live, before more iteration on the upload path. I think it splits cleanly in two, and neither half belongs in the Harbor example. 1. Getting records to durable storage is skycap's job. skycap already writes each trajectory once, when it ends, off the event loop. The semantics we want for the upload (async, fail open, bounded queue, timeouts, a deadline at shutdown) are exactly a record writer's. A remote store is just another destination for the same records. It also removes the multi-node problem at the source: today records land on each server's node-local disk, so this PR has to fetch them back over HTTP through the trainer and re-upload the bytes. If skycap writes to shared or remote storage (e.g. 2. What goes to W&B is the trainer's job. "The trajectories of training step N, their phase, which were trained and which superseded" is SkyRL knowledge. A callback can log that per step through With that split, most of this PR's fetch/re-upload/retry/timeout machinery goes away. I'm opening the two pieces as separate PRs and will link them here. Happy to fold anything from this PR's design into them (the |
|
Following up on the comment above, the two pieces are up:
Together they replace this PR's upload path. Reviews and pushback welcome there. |
SkycapRecordIndex, a trainer callback, logs one version per step of
`skycap-records-<phase>-<run id>`, aliased `<phase>-step-N` and `latest`:
a step.json with every attempt the generator opened (trained or
superseded, and its record location from FinishResult.record), plus a
checksum-free W&B reference per mirrored document. No record bytes are
uploaded. The index is built on the trainer's thread; the W&B calls run on
a background thread and fail open (bounded queue, per-call timeout without
retry, bounded retries before W&B accepts a version, shutdown deadline).
The generator logs its trajectories into a RecordLog per phase. The
trainer passes the batch's trajectory_ids to on_step_end, so the index can
tell which trajectories trained. skycap.record_mirror is passed through to
the servers, and skycap.wandb.{enabled,phases} configures the index.
Builds on #2333's design (SkycapWandbConfig, per-phase RecordLog, the
trained/superseded marking, eval opt-in).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com>
SkycapRecordIndex, a trainer callback, logs one version per step of
`skycap-records-<phase>-<run id>`, aliased `<phase>-step-N` and `latest`:
a step.json with every attempt the generator opened (trained or
superseded, and its record location from FinishResult.record), plus a
checksum-free W&B reference per mirrored document. No record bytes are
uploaded. The index is built on the trainer's thread; the W&B calls run on
a background thread and fail open (bounded queue, per-call timeout without
retry, bounded retries before W&B accepts a version, shutdown deadline).
The generator logs its trajectories into a RecordLog per phase. The
trainer passes the batch's trajectory_ids to on_step_end, so the index can
tell which trajectories trained. skycap.record_mirror is passed through to
the servers, and skycap.wandb.{enabled,phases} configures the index.
Builds on #2333's design (SkycapWandbConfig, per-phase RecordLog, the
trained/superseded marking, eval opt-in).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com>
SkycapRecordIndex, a trainer callback, logs one version per step of
`skycap-records-<phase>-<run id>`, aliased `<phase>-step-N` and `latest`:
a step.json with every attempt the generator opened (trained or
superseded, and its record location from FinishResult.record), plus a
checksum-free W&B reference per mirrored document. No record bytes are
uploaded. The index is built on the trainer's thread; the W&B calls run on
a background thread and fail open (bounded queue, per-call timeout without
retry, bounded retries before W&B accepts a version, shutdown deadline).
The generator logs its trajectories into a RecordLog per phase. The
trainer passes the batch's trajectory_ids to on_step_end, so the index can
tell which trajectories trained. skycap.record_mirror is passed through to
the servers, and skycap.wandb.{enabled,phases} configures the index.
Builds on #2333's design (SkycapWandbConfig, per-phase RecordLog, the
trained/superseded marking, eval opt-in).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com>
SkycapRecordIndex, a trainer callback, logs one version per step of
`skycap-records-<phase>-<run id>`, aliased `<phase>-step-N` and `latest`:
a step.json with every attempt the generator opened (trained or
superseded, and its record location from FinishResult.record), plus a
checksum-free W&B reference per mirrored document. No record bytes are
uploaded. The index is built on the trainer's thread; the W&B calls run on
a background thread and fail open (bounded queue, per-call timeout without
retry, bounded retries before W&B accepts a version, shutdown deadline).
The generator logs its trajectories into a RecordLog per phase. The
trainer passes the batch's trajectory_ids to on_step_end, so the index can
tell which trajectories trained. skycap.record_mirror is passed through to
the servers, and skycap.wandb.{enabled,phases} configures the index.
Builds on #2333's design (SkycapWandbConfig, per-phase RecordLog, the
trained/superseded marking, eval opt-in).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com>
**TLDR:** each training step's skycap records are indexed in W&B: one artifact version per step listing which trajectories made up the step, and how each one ended up. Mirrored records (#2350) appear as W&B *references*, so no record bytes are uploaded. It replaces #2333's upload path, and its design (config shape, per-phase log, eval opt-in) comes from that PR. --------- Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Adds
SkycapUploads, a trainer callback that uploads skycap documents to W&B whentrainer.logger=wandb. It fetches each document from its skycap server and uploads on a background thread;on_train_endwaits for it.Train records upload once per step to
skycap-records-train-<run id>, aliasedstep-Nandlatest.skycap.wandb.phasesadds eval, which goes to its ownskycap-records-eval-<run id>artifact.skycap.wandb.enabled=falseturns uploads off.Each version holds a
step.jsonthat marks every attempt trained or dropped (retries, dynamic sampling). The metadata counts uploaded and missing documents. For this,CallbackInputgainstrajectory_idsaton_step_end.The fully-async trainer fires no callbacks, so it uploads nothing.
run_codecontests_modal.shalso takes the Modal credentials from the environment, and its default run names carry a UTC timestamp.Tested:
tests/integrations/harbor_skycapandtests/train/test_rl_callbacks.pypass on 8xB200 (21 passed).🤖 Generated with Claude Code