Your AI. Your Device. Your Silicon.
Latest Product Updates
Gemma 4 (E2B + E4B)
Apache 2.0, 128K context. New SoTA local class. Support for Gemma 4n with MTP speculative decoding — up to 2× faster on mobile GPU.
NPU on Snapdragon
QNN delegate via Play Feature Delivery. SoC-aware backend selection (QNN → GPU → CPU) for lower battery drain and NPU speeds.
MLX on Apple Silicon
Real inference on macOS & iOS 18+ A17 Pro+. Metal-native execution with 1-bit affine quantization, bringing 7B models down to ~1.75 GB.
On-Device AI Agents
Plan-and-execute agents with task memory, scheduled cron runs, composable skills, and native mobile tools (clipboard, contacts, files).
Why FluentAI?
A fully private AI agent platform that runs models natively on your device without monthly subscriptions or data leaks.
Privacy First
Conversations stay entirely on-device. Zero telemetry, zero tracking, and zero data collection. No cloud server required.
100+ AI Models
Run Llama 3, Gemma 4, DeepSeek, Mistral, Phi, or Qwen locally. Or optionally connect to Claude, GPT-4, and Gemini in the cloud using your own API keys.
Multi-Runtime Engine
Three inference backends (Fllama/GGUF, LiteRT/Android NPU, MLX/Apple Silicon) built-in. The app auto-selects the fastest for your device.
Knowledge Bases
Upload PDFs and text files to query your data locally. Fully private Retrieval-Augmented Generation (RAG) with local semantic search.
Tool Calling & MCP
Built-in tools for search, calculations, and memory. Full Model Context Protocol (MCP) support connects your AI to GitHub, Slack, Notion, and more.
Voice Chat
Speak with your AI naturally using 5 distinct conversation modes: Normal, Interview, Learning, Storytelling, and Translation.
Local OpenAI API Server
FluentAI serves /v1/chat/completions directly on your local network. Other apps can use FluentAI as their offline model backend.
Hugging Face Browser
Search 10,000+ GGUF models directly in-app. Features memory-fitness badges to prevent your phone or tablet from running out of RAM (OOM).
One App. Three Runtimes.
FluentAI embeds multiple inference frameworks to deliver the fastest performance regardless of your hardware configuration.
FllamaRuntime
GGUF · llama.cpp · All Platforms
- Gemma 4 architecture backport (MoE 128 experts, ISWA dual-cache)
- KleidiAI v1.23.0 optimizations (SME2 + Q4_K paths)
- KV cache TQ4/TQ3 quantization for memory efficiency
- 16 KB page alignment support for Android 15+
LiteRTRuntime
Android · NPU / GPU · LiteRT-LM 0.10
- Snapdragon NPU acceleration via QNN delegate
- SoC-aware backend selection: QNN → GPU → CPU
- Play Feature Delivery for modular, bloat-free installs
- MTP speculative decoding for 1.5–2× faster generation
MlxRuntime
macOS · iOS 18+ A17 Pro+ · Apple Silicon
- Native Apple MLX inference on M-series chips and iOS
- 1-bit affine quantization runs 7B models in ~1.75 GB RAM
- Metal-native execution with zero Rosetta overhead
- Multi-file parallel downloads direct from Hugging Face
Supported Platforms & Downloads
Google Play Store
Full hardware NPU acceleration, GPU/CPU execution, and offline local model downloads.
macOS, Windows & Linux
Run MLX models on Apple Silicon macOS, or load local GGUF models via Fllama on Windows and Linux.
Pricing Plans
Choose the plan that suits you best. Support independent open-source AI development.
Free Tier
Perfect for offline, private local AI execution.
- 100+ local AI models support
- Unlimited offline chat conversations
- Full voice chat with 5 modes
- Knowledge bases (Private RAG)
- Tool Calling & MCP support
Premium Upgrade
Unlock customization options and support the product.
- Everything in Free Tier
- 100% Ad-free experience
- 9 premium visual themes
- Private cloud sync dashboard
- Advanced model temperature/settings
- PDF conversation exports & analytics
Frequently Asked Questions
Is FluentAI really free?
Yes! Running local models on your device is completely free and unlimited. If you decide to connect to cloud providers like Claude, GPT-4, or Gemini, you will just need to supply your own API keys.
Does it work completely offline?
Absolutely. Once you download your chosen GGUF or LiteRT models, all conversation, processing, and inference are executed entirely offline on your device, requiring no internet connection whatsoever.
Which AI models can I run?
You can run over 100 models locally including Llama 3, Gemma 4 (E2B/E4B), DeepSeek, Mistral, Phi, and Qwen via GGUF, LiteRT, or MLX formats. You can also import any GGUF model directly by pasting its Hugging Face URL.
How does the NPU acceleration work?
On compatible Android devices with Snapdragon SoCs, FluentAI uses the LiteRT QNN delegate to route tensor math directly to the Snapdragon NPU. This yields up to 4× speedups over standard CPU execution while using significantly less battery.
What platforms are supported?
Android is available now on Google Play. macOS and iOS 18+ (A17 Pro+) versions support real Apple Silicon MLX inference. Desktop versions for Windows and Linux are also available for download.
Start Chatting Privately Today
No credit card, no sign-up, no monthly fees. Just pure on-device AI running with hardware acceleration. Take back your privacy.
