Favicon of LazyLLM

LazyLLM

Python AI agent framework for self-hosted applications, with local inference, cloud model APIs, RAG and fine-tuning. Open source under Apache 2.0.

Screenshot of LazyLLM website

LazyLLM is a Python framework for developers building multi-agent applications with locally deployed models or cloud services. It combines application assembly with data preparation, evaluation and model fine-tuning, so developers can test a prototype and improve the parts that perform poorly. It's open source under Apache 2.0.

Its workflow modules support branching, loops and parallel tasks. Developers can combine these with model and retrieval components to build chatbots or document question-answering applications. Built-in browser interfaces cover chat and document management, while MCP integration supports agent task planning and tool use.

Local inference backends include vLLM, LMDeploy and LightLLM, with Infinity for embeddings. For offline fine-tuning, it supports LLaMA-Factory, Alpaca-LoRA and Collie. A common interface for local models and online providers lets developers compare model choices within the same application. Cloud connections include OpenAI, SenseNova, Kimi, Qwen and Zhipu; these require provider accounts and API keys, and send requests to the provider. Local model services run on your own hardware.

The framework also includes document parsing, retrieval and reranking for RAG, plus speech generation and text-to-image capabilities. It runs on development machines, bare-metal servers and Slurm clusters, and has a Docker image. Applications can be packaged for deployment with Kubernetes load balancing and fault tolerance.

Similar to LazyLLM