Files
LTX-2/packages/ltx-trainer/configs/README.md
indigo e03cc62548 Add SCAIL-2 training integration (Phase 3, ltx-trainer FlexibleStrategy)
Wire SCAIL-2 driving + in-context mask conditioning into the trainer via the
unified FlexibleStrategy, so the widened patchify_proj (Phase 2) can be trained.

- flexible.py: new DrivingConditionConfig (driving-latent concat with a RoPE
  width offset ΔW) and MaskChannelsConditionConfig (semantic masks -> per-token
  channels), added to the condition union and get_data_sources. Driving is
  prepended (cond-first, target stays at the tail for loss slicing); mask
  channels are written onto the driving tokens via Modality.cond_channels,
  reusing ltx-core encode_mask_channels. The noisy target keeps a zero mask.
- model_loader.load_transformer gains mask_conditioning_channels, widening the
  video patchify_proj with zero-init columns via a new live-module helper
  (widen_module_patchify_proj_for_mask_channels). ModelConfig exposes the field.
- trainer unfreezes patchify_proj in LoRA mode when mask channels are active
  (the new input columns are new base params LoRA cannot reach).
- configs/scail_animation_lora.yaml plus README / training-modes table rows.

Verified on CPU (verify_phase3_trainer.py): config round-trips, prepare_training
_inputs builds cond_channels [B,T,56] with the mask on the driving tokens, a tiny
widened model forwards and compute_loss returns a finite [B] loss, and the widen
helper is output-preserving at zero init. Real training needs Linux+GPU+checkpoint;
dataset preprocessing (driving latents + semantic masks) and validation-runner
wiring are left for later.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 10:09:40 +08:00

2.9 KiB

Training Configs

Example training configurations for the LTX-2 trainer. Each file is a ready-to-run config for one training mode, expressed through the unified flexible strategy (name: "flexible"). Pick the one closest to your use case and adjust paths, dataset, and hyperparameters.

📖 For more information about using each training mode, see Training Modes Guide.

Training Modes

Mode Video Audio Conditions Config
T2V Generated Generated t2v_lora.yaml, t2v_lora_low_vram.yaml (low VRAM)
I2V Generated Generated first_frame i2v_lora.yaml
Video Extension Generated Generated prefix/suffix video_extend_lora.yaml (forward), video_suffix_lora.yaml (backward)
V2V IC-LoRA Generated reference v2v_ic_lora.yaml
A2V Generated Frozen a2v_lora.yaml
V2A (Foley) Frozen Generated v2a_lora.yaml
Video Inpainting Generated mask video_inpainting_lora.yaml
Video Outpainting Generated spatial_crop video_outpainting_lora.yaml
T2A Generated t2a_lora.yaml
Audio Extension Generated prefix/suffix audio_extend_lora.yaml (forward), audio_suffix_lora.yaml (backward)
Audio Inpainting Generated mask audio_inpainting_lora.yaml
A2A IC-LoRA Generated reference a2a_ic_lora.yaml
AV2AV IC-LoRA Generated Generated reference (both) av2av_ic_lora.yaml
SCAIL Animation Generated driving + mask_channels scail_animation_lora.yaml

The accelerate/ directory holds the Accelerate launch configs (FSDP, DDP) for multi-GPU training.

SCAIL Animation (SCAIL-2 character animation) also sets model.mask_conditioning_channels: 56 to widen the video patchify_proj for the in-context mask channels; in LoRA mode that projection is unfrozen so the new columns train. See Training Modes Guide.