fsdp
Here are 72 public repositories matching this topic...
A scalable, agentic-first, and HuggingFace-native RL framework for research (9k lines).
-
Updated
Oct 8, 2026 - Python
Best practices & guides on how to write distributed pytorch training code
-
Updated
Oct 22, 2025 - Python
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
-
Updated
Oct 9, 2026 - Python
Find why PyTorch training is slow, compare runs, and catch performance regressions in CI.
-
Updated
Oct 9, 2026 - Python
Educational distributed training and inference library for heterogeneous local hardware: Mac minis, Raspberry Pis and GPUs over plain Python sockets. FSDP, GRPO, parameter-server, expert and data parallelism.
-
Updated
Oct 7, 2026 - Python
Write once, run anywhere; ezpz 🍋
-
Updated
Sep 30, 2026 - Python
Research platform for model training, evaluation, and experimentation across architectures, benchmarks, and recipes.
-
Updated
Oct 1, 2026 - Python
Fast and easy distributed model training examples.
-
Updated
Nov 26, 2024 - Python
Forge kernels — fused Triton kernels for faster, leaner LLM fine-tuning. One-call patching into Hugging Face models via forge.patch(model), with FSDP2 multi-GPU support and a growing set of kernels and architectures. Every speedup backed by a committed benchmark. Apache-2.0.
-
Updated
Aug 15, 2026 - Python
A script for training the ConvNextV2 on CIFAR10 dataset using the FSDP technique for a distributed training scheme.
-
Updated
Dec 11, 2023 - Python
Simple and efficient implementation of 671B DeepSeek V3 that trainable with FSDP+EP and minimal requirement of 256x A100/H100, targeted for HuggingFace ecosystem
-
Updated
Jan 15, 2026 - Python
Minimal yet high performant code for pretraining llms. Attempts to implement some SOTA features. Implements training through: Deepspeed, Megatron-LM, and FSDP. WIP
-
Updated
Feb 6, 2024 - Python
Framework, Model & Kernel Optimizations for Distributed Deep Learning - Data Hack Summit
-
Updated
Aug 1, 2023 - Python
Distributed fine-tuning on Kubernetes that survives mid-run pod eviction and resumes bit-identically.
-
Updated
Aug 8, 2026 - Python
Add this topic to your repo
To associate your repository with the fsdp topic, visit your repo's landing page and select "manage topics."