Multi-node PyTorch DDP fine-tuning of a causal LM on Nebius GPU Kubernetes, with SkyPilot workload orchestration — a 2-node torchrun job with verified NCCL collectives.
-
Updated
Jun 26, 2026 - Python
Multi-node PyTorch DDP fine-tuning of a causal LM on Nebius GPU Kubernetes, with SkyPilot workload orchestration — a 2-node torchrun job with verified NCCL collectives.
This is an end-to-end distributed deep learning orchestrator designed to scale model training across multi-node clusters. It simplifies Distributed Data Parallel (DDP), networking, SSH orchestration, and telemetry by unifying them into a seamless workflow with a centralized, real-time web dashboard for monitoring, control, and performance insights.
Add a description, image, and links to the torchrun topic page so that developers can more easily learn about it.
To associate your repository with the torchrun topic, visit your repo's landing page and select "manage topics."