FlashAttention is a GPU attention library for developers training or running transformer models on their own hardware or servers. It computes exact attention with less memory use than standard PyTorch attention, helping models handle longer sequences. It's free and open source under the BSD-3-Clause license.
Its attention memory use grows linearly with sequence length, while standard attention's grows quadratically. Speed gains depend on the GPU and workload, so benchmark results aren't a fixed promise for every model. The library includes forward and backward computation for training, plus KV-cache support for incremental text generation.
FlashAttention works with PyTorch on Linux using NVIDIA CUDA or AMD ROCm. Supported NVIDIA hardware includes Ampere, Ada and Hopper GPUs, such as the RTX 3090, RTX 4090, A100 and H100. A CuTeDSL implementation targets Hopper and Blackwell hardware, including B200. AMD support covers Instinct and RDNA GPUs through Composable Kernel and Triton backends. Windows builds require further testing.
Model developers get causal and sliding-window attention, multi-query and grouped-query attention, and support for positional methods such as rotary embeddings and ALiBi. The AMD Triton backend also supports paged attention and variable sequence lengths. Beyond attention, the project includes a GPT implementation, training scripts for GPT-2 and GPT-3, and optimized layers for MLP, LayerNorm and cross-entropy loss.
Claim this page and we'll verify you by hand. FlashAttention gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find FlashAttention?Promote it
Something wrong or outdated on this page?
2.8KUpdated 9 months agoApache-2.0
Windows · Linux#Batch processing#ONNX#Voice activity detection
openWakeWord is a Python library for developers building voice interfaces that listen locally for a chosen word or phrase. It includes English models for triggers such as "hey jarvis" and "alexa", plus phrases for weather and timers. The code uses Apache 2.0. Included pretrained models use CC-BY-NC-SA-4.0, which restricts commercial use.
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
37.1KUpdated 1 day agoApache-2.0
iOS · Android · Web
MediaPipe is an open-source toolkit for developers adding on-device AI to applications on Android, iOS, the web, desktop and edge devices. It pairs pretrained models with APIs for specific tasks, so developers can use existing solutions or customize them for their applications. The project uses the Apache 2.0 license.
7.7KUpdated 2 years agoMIT
Docker · Web#Multilingual
2.2KUpdated 15 hours agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.