
Jamba is AI21's language model family for teams building AI applications on their own servers. The documented Large 1.7 model combines Mamba state-space models with Transformer attention to process long context efficiently. A 256K-token context window makes it relevant for work that depends on lengthy documents, such as investment research, due diligence and reviewing procurement responses.
The model focuses on answers grounded in the material you provide and on following instructions about how to respond. Those traits matter for document-based assistants and customer chatbots that need to stay within supplied context or follow a brand's writing requirements. Other stated uses include drafting medical reports, product descriptions and grant applications, as well as personalized tutoring.
Hardware needs are substantial. The Large model requires quantization to fit on a server with eight GPUs, each with 80GB of memory. It works with vLLM and Hugging Face Transformers; AI21's ExpertsInt8 quantization supports prompts up to 220K tokens on that hardware with the default vLLM configuration. Private deployment runs inference on your own infrastructure. Fine-tuning options include LoRA, qLoRA and full fine-tuning.
The Jamba Open Model License allows research and commercial use under its terms. Access to the model files on Hugging Face requires an account, acceptance of the access conditions and sharing contact information. Supported languages include English, Spanish, French, Portuguese, Italian, Dutch, German, Arabic and Hebrew.
Claim this page and we'll verify you by hand. Jamba gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Jamba?Promote it
Something wrong or outdated on this page?
107Updated 1 year ago
#GGUF#Hugging Face integration#Multilingual
EXAONE 4.0 is a family of language models from LG AI Research that combines general language tasks and complex problem solving in the same model. It's aimed at developers building multilingual AI applications, including on-device apps and agents that use tools. It supports English, Korean and Spanish.
127Updated 12 months ago
macOS#GGUF#Hugging Face integration#llama.cpp backend
Falcon-H1 is a family of language models from the Technology Innovation Institute for developers building local AI applications or serving models on their own hardware. Its hybrid design combines attention with Mamba2 state space models to reduce memory use and speed up inference. Model sizes cover smaller devices and larger server deployments.
huggingface.coCoding Models
#Hugging Face integration#LoRA#Multilingual
27.7KUpdated 9 months ago
#Batch processing#GGUF#Hugging Face integration
7.8KUpdated 2 years agoApache-2.0
Docker#Hugging Face integration#llama.cpp backend#Multilingual
huggingface.coOpen-Weight LLMs
#Guardrails#Hugging Face integration#Multilingual
GLM-4.5 is an open-source language model for developers building AI agents and coding tools on their own servers. It combines reasoning with tool calling and offers a choice between thinking mode for complex tasks and non-thinking mode for direct responses. The MIT license permits commercial use and modification.
Qwen3 is a family of language models from Alibaba Cloud’s Qwen team for people who want to run models locally or on their own servers. It spans smaller and larger dense models as well as mixture-of-experts models. The weights are publicly available.
Yi is a family of open-weight language models from 01.AI for developers, researchers, and businesses that want to run English and Chinese models on their own hardware. It includes models for conversation and base models for fine-tuning, with license terms that depend on the release. Yi-1.5 code and weights use Apache 2.0; the original Yi releases have separate community terms.
Command A is an open weights language model from Cohere and Cohere Labs for researchers and developers building self-hosted chatbots, document assistants and AI agents. Its focus is business tasks that combine multilingual text, supplied documents and external tools. You can run the model on your own hardware; Cohere also offers hosted chat through a playground and Hugging Face Space.