Repository navigation
Actions: google/tunix
Actions
Showing runs from all workflows
2,500+ workflow runs
2,500+ workflow runs
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
Tunix Package Tests
#11075:
Pull request #2729
synchronize
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
Check the Documentation Build
#7344:
Pull request #2729
synchronize
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
GitHub Actions Scan
#3994:
Pull request #2729
synchronize
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
Tunix Package Tests
#11071:
Pull request #2729
synchronize
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
Check the Documentation Build
#7340:
Pull request #2729
synchronize
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
GitHub Actions Scan
#3990:
Pull request #2729
synchronize
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
Tunix Package Tests
#11070:
Pull request #2729
opened
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
GitHub Actions Scan
#3989:
Pull request #2729
opened
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
Auto-assign issues and PRs
#1630:
Pull request #2729
opened
by
copybara-service
Bot
PeftTrainer.grad_accumulator to persistent mode (allocate_grads=True) and clear the JIT cache on first fwd_bwd() invocation so split fwd_bwd() + update() calls under nnx.cached_partial (cache_nnx_graph=True) and dynamic sequence-packing microsteps (gradient_accumulation_steps == 1) have a pre-allocated, sharded gradient buffer across the JIT boundary, while preserving the zero-allocation fused train_step() fast path for train().
Check the Documentation Build
#7339:
Pull request #2729
opened
by
copybara-service
Bot