Favicon of Falcon

Falcon

A local LLM family for chat, coding and multilingual tasks, with GGUF and Hugging Face formats for CPU or GPU use and support for llama.cpp and MLX.

Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.

The models handle instruction following, coding, maths and translation. Their multilingual training includes Arabic, English, Hindi, Japanese and Chinese, among other languages. Context windows reach up to 256K tokens for long documents and extended conversations, though the maximum depends on the model.

GGUF and Hugging Face formats give users a choice of runtimes. Falcon-H1 works with llama.cpp and MLX for local use, as well as Hugging Face Transformers, vLLM and SGLang. It can run on CPUs or GPUs, and the project demonstrates a quantized model on a MacBook with an M4 Max chip. SGLang deployments require a compatible NVIDIA GPU and can expose an OpenAI-compatible API.

Local inference runs on your hardware; the hosted chat interface and Hugging Face demo are separate online options. OpenWebUI can provide a browser chat interface for a locally served model. For teams adapting the models to their own tasks, fine-tuning support includes Llama-Factory, Unsloth and OUMI.

Falcon-H1 model weights use the Falcon-LLM License, whose conditions are separate from the Apache-licensed inference code.

Similar to Falcon