nanochat is an MIT-licensed toolkit for training your own LLM and chatting with it on hardware you control. It's aimed at researchers and developers who want to study or modify the full training pipeline, with a small Python codebase built on PyTorch.
The pipeline covers tokenizer training, pretraining, supervised finetuning, reinforcement learning, evaluation and inference. A command-line chat interface lets you talk to the resulting model, and the model can execute Python code as a tool. Evaluation includes science, math and coding tasks, alongside benchmarks for inference speed and memory use.
Its distinctive approach is to tie model size and training settings to transformer depth. Choosing a smaller or larger model automatically adjusts its width, attention heads and training hyperparameters. That makes it useful for comparing models at different scales without managing a large collection of separate settings. The repository includes experiments for studying scaling laws and measuring training time against GPT-2 capability on the DCLM CORE benchmark.
Training runs on a single GPU node, either on your own hardware or a rented server. The reference setup uses NVIDIA H100 GPUs, and A100 GPUs also work. A single GPU can run the pipeline more slowly; GPUs with less memory need adjusted batch sizes. CPU and Apple Silicon examples train much smaller models with limited capabilities. Research runs can use Weights & Biases to track training metrics.
Claim this page and we'll verify you by hand. nanochat gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find nanochat?Promote it
Something wrong or outdated on this page?
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.