scripts/process_char_masks.py turns per-sample character label-map videos/images
(integer pixel labels: 0 = environment, 1..K = characters -> binding slots) into
the pixel-space semantic-mask tensors the SCAIL training/inference path consumes:
{"mask": [K+1, F_pix, H_pix, W_pix]} (ch0 = environment switch, ch1..K = slots).
- Aligns to the target video's latent grid read from the saved latent metadata
(F_pix=(F-1)*8+1, H*32, W*32), so char_masks/ lines up file-for-file with
latents/ / driving_latents/ for PrecomputedDataset.
- Nearest-neighbour resize so integer labels are never blended; labels > K are
dropped with a warning; ch0 filled uniformly with --environment-switch.
- Reuses process_videos.py helpers (naming, atomic save, VAE factors) and matches
its typer CLI conventions.
Verified on CPU: a synthetic 2-character label map (plus an out-of-range id)
produces mask (7,17,128,128) with ch0 uniform, slots placed correctly, id>K
dropped, and feeds encode_mask_channels to the 8*(K+1)=56 channels. README +
docs/tasks.md 3.7 updated (upstream label-map generation via SAM/tracking is
dataset-specific and still out of scope).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
LTX-2 Trainer
This package provides tools and scripts for training and fine-tuning Lightricks' LTX-2 audio-video generation model. It supports LoRA training, full fine-tuning, and a flexible conditioning framework covering text-to-video, text-to-audio, image-to-video, video extension, audio extension, video inpainting, audio inpainting, video outpainting, IC-LoRA for video, audio, and joint audio-video references, audio-to-video, and video-to-audio.
📖 Documentation
All detailed guides and technical documentation are in the docs directory:
- ⚡ Quick Start Guide
- 🎬 Dataset Preparation
- 🛠️ Training Modes
- ⚙️ Configuration Reference
- 🚀 Training Guide
- 🧪 Inference Guide
- 🔧 Utility Scripts
- 🧩 Custom Training Strategies
- 📚 LTX-Core Documentation
- 🛡️ Troubleshooting Guide
🤖 Agent-Assisted Training
Use the train-model repository skill for an end-to-end guided run:
it probes your data and hardware, chooses the matching training mode, prepares/preprocesses the dataset, launches
training, and monitors the job while using the docs above as the source of truth.
🔧 Requirements
- LTX-2 Model Checkpoint - Local
.safetensorsfile - Gemma Text Encoder - Local Gemma model directory (required for LTX-2)
- Linux with CUDA - CUDA 13+ recommended for optimal performance
- Nvidia GPU with 80GB+ VRAM - Recommended for the standard config. For GPUs with 32GB VRAM (e.g., RTX 5090), use the low VRAM config which enables INT8 quantization and other memory optimizations
🤝 Contributing
We welcome contributions from the community! Here's how you can help:
- Share Your Work: If you've trained interesting LoRAs or achieved cool results, please share them with the community.
- Report Issues: Found a bug or have a suggestion? Open an issue on GitHub.
- Submit PRs: Help improve the codebase with bug fixes or general improvements.
- Feature Requests: Have ideas for new features? Let us know through GitHub issues.
💬 Join the Community
Have questions, want to share your results, or need real-time help?
Join our community Discord server to connect with other users and the development team!
- Get troubleshooting help
- Share your training results and workflows
- Collaborate on new ideas and features
- Stay up to date with announcements and updates
We look forward to seeing you there!
Happy training! 🎉