Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
The models handle instruction following, coding, maths and translation. Their multilingual training includes Arabic, English, Hindi, Japanese and Chinese, among other languages. Context windows reach up to 256K tokens for long documents and extended conversations, though the maximum depends on the model.
GGUF and Hugging Face formats give users a choice of runtimes. Falcon-H1 works with llama.cpp and MLX for local use, as well as Hugging Face Transformers, vLLM and SGLang. It can run on CPUs or GPUs, and the project demonstrates a quantized model on a MacBook with an M4 Max chip. SGLang deployments require a compatible NVIDIA GPU and can expose an OpenAI-compatible API.
Local inference runs on your hardware; the hosted chat interface and Hugging Face demo are separate online options. OpenWebUI can provide a browser chat interface for a locally served model. For teams adapting the models to their own tasks, fine-tuning support includes Llama-Factory, Unsloth and OUMI.
Falcon-H1 model weights use the Falcon-LLM License, whose conditions are separate from the Apache-licensed inference code.
Claim this page and we'll verify you by hand. Falcon gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Falcon?Promote it
Something wrong or outdated on this page?
20.4KUpdated 2 months agoApache-2.0
macOS · Linux#Hugging Face integration#LM Studio integration#Ollama integration
gpt-oss is a pair of OpenAI reasoning models for developers who want to run a local LLM or host one on their own server. The models are open weight and licensed under Apache 2.0. OpenAI also has a hosted browser demo, separate from running the models on your hardware.
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
huggingface.coOpen-Weight LLMs
#Hugging Face integration#LoRA#Multilingual
Jamba is AI21's language model family for teams building AI applications on their own servers. The documented Large 1.7 model combines Mamba state-space models with Transformer attention to process long context efficiently. A 256K-token context window makes it relevant for work that depends on lengthy documents, such as investment research, due diligence and reviewing procurement responses.
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
27.7KUpdated 9 months ago
#Batch processing#GGUF#Hugging Face integration
Qwen3 is a family of language models from Alibaba Cloud’s Qwen team for people who want to run models locally or on their own servers. It spans smaller and larger dense models as well as mixture-of-experts models. The weights are publicly available.