feat(data): shard-backed dataset, H5-to-shard converter, planner_shards package - #416
Closed
Dionysus326 wants to merge 1 commit into
Closed
Dionysus326 wants to merge 1 commit into
Dionysus326 wants to merge 1 commit into
Conversation
…ds package Add a tar-shard data path beside the HDF5 loader: - packages/planner_shards: the versioned shard dataset library (packer, manifests, key-sets, DDP-safe loader) as a workspace package; a copy of the tier4-main data_pipeline/shard_* modules with imports renamed, since both trees use the diffusion_planner package name. - ShardPlannerDataset / build_shard_dataloader apply the PlannerDataset transforms to frames read from shards; rank and world size come from the torchrun environment. - configs/train/dataloader/shards.yaml mirrors default.yaml's transforms. - train.py keeps a shard-backed loader out of accelerator.prepare (it is already sharded per rank) and moves batches to the device itself. - convert_h5_to_shards.py stages the frames of a Parquet index in chunks, packs them with the unchanged packer, writes a key-set, scrubs, and spot-checks members bit-exact against the H5 source.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a tar-shard data path to the new-architecture training tree, next to the existing HDF5 loader:
packages/planner_shards— the versioned shard dataset library (packer, manifests, key-sets, DDP-safe loader) as a uv workspace package. It is a copy of thedata_pipeline/shard_*modules from thetier4-maintree (feat(data_pipeline): versioned shard dataset (WebDataset-style tars + parquet manifest) with DDP loader #390) with imports renamed, because both trees use the package namediffusion_planner. Deliberate temporary duplication: neither PR can depend on the other while both are drafts; once one lands, the other switches to a dependency.ShardPlannerDataset/build_shard_dataloader(diffusion_planner.data) — iterates one rank's share of a packed dataset and applies the same frame transforms asPlannerDataset. Rank and world size come from the torchrun environment.configs/train/dataloader/shards.yaml— same transform stack asdefault.yaml; select withtrain/dataloader@dataloader=shards.scripts/train/train.py— when the loader wrapsShardPlannerDataset, it is not passed throughaccelerator.prepare(it is already sharded per rank) and batches are moved to the device in the loop. The HDF5 path is unchanged.scripts/dataset/convert_h5_to_shards.py(diffusion_planner.data.h5_to_shards) — converts the frames of a Parquet index into shards: one member per frame, one partition per rosbag, staged in chunks through a scratch directory and packed by the unchanged packer; writes a key-set, scrubs, and spot-checks members bit-exact against the H5 source.Testing
packages/diffusion_planner/tests/data/test_shard_planner_dataset.py— two ranks cover every frame exactly once, bit-exact; transforms apply in order on writable arrays; loader length matches the plan; loader settings must match the plan.packages/diffusion_planner/tests/data/test_h5_to_shards.py— chunked conversion of a synthetic H5 layout, key-set size, intermediate and final versions, bit-exact verification, staging cleanup.packages/planner_shards/tests— the library's own suite (108 tests) with imports renamed.train/traincomposed withtrain/dataloader@dataloader=shardsinstantiates and yields batches.tests/datasuite still passes.Measured on an internal 8-GPU box: whole training epochs over identical frames were 1.5x faster through shards than through the HDF5 loader at its default worker count; details are internal. Not a convergence comparison.
Notes for review
planner_shardsis not yet in the pyright include list; the copied library predates type-checking and will be cleaned up when it becomes a shared dependency..npzbefore packing (the packer's existing input contract). A direct H5-to-member writer is a follow-up.