claudeers.
// MCP Servers

Box

The most advanced, fully offline client-side AI suite on Android today.

// MCP Servers[ cli ][ api ][ desktop ][ web ][ mobile ][ claude ]#claude#android#android-ai-app#artificial-intelligence#box#gguf#linux#linux-ai#mcp-servers◷ NOASSERTION$open-sourceupdated about 1 month ago
Actively maintained
100/100
last commit 15 days ago
last release 15 days ago
releases 34
open issues 1
// star history

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up Box (release-binary project) into my current project.
Found on https://claudeers.com/box
Repo: https://github.com/jegly/Box
Homepage/docs: https://jegly.xyz
Detected install method: release-binary → inspect the README
Category: mcp-servers. Platforms: cli, api, desktop, web, mobile.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
active; community-verified: false. Confirm the source before running anything.
// or install directly (release-binary)

Grab the latest release asset from GitHub.

# download a build from https://github.com/jegly/Box/releases
// or clone
git clone https://github.com/jegly/Box

// compatibility

Platformscli, api, desktop, web, mobile
Operating systems—
AI compatibilityclaude
LicenseNOASSERTION
Pricingopen-source
LanguageKotlin

Get your FREE $2.50 API credits to access TickAtlas financial data ↗

Box Header

⭐️ If this project helped you, please star it — it helps others find it.

We've hit 28K downloads! Thank you to everyone for supporting Box.

Note: If you're using a custom ROM (LineageOS, GrapheneOS, CalyxOS), download the custom-rom-support APK from the latest release instead.

Install via Obtainium

  1. Open Obtainium on your phone
  2. Tap the + button
  3. Paste this repo URL:
    https://github.com/jegly/Box
  4. Tap Add

Recommended for most users: Main version

Which version should I install?

VersionFor
MainStock Android (Pixel, Samsung, etc.)
Custom ROMGrapheneOS, LineageOS, CalyxOS — no Google services
  • The in-app updater is also available in Settings

    Setup steps

    1. Tap the badge for your version above — this opens Obtainium with the repo pre-filled
    2. Under APK filter regex, enter one of the following:
      • Main: Main
      • Custom ROM: custom-rom-support
    3. Tap Add — Obtainium will find the latest release and install it
    4. Future updates will be detected automatically

    Note: As of v2.0.0, the in-app App version matches the Box release version (2.0.0) — the earlier mismatch with the upstream Google AI Edge Gallery build number (which showed 1.0.15) is fixed (#67). Box releases are tracked via GitHub tags. Use Settings → Check for updates to see if a newer Box release is available.

Box is a security-hardened, feature rich fork of Google AI Edge Gallery — with on-device image generation (Bonsai Image 4B, FLUX.2 klein & Z-Image Turbo diffusion), Box Assist (spoken camera assistance for blind and low-vision users), AI image upscaling, face recognition, photo erase/inpainting, music & sound generation, voice mode (speech-to-speech AI chat), voice input, multilingual text-to-speech, document analysis and Q&A, vision AI, full GPU and Snapdragon/Tensor/MediaTek NPU acceleration, a hardened security posture (biometric lock, encrypted chat history, tap-jacking protection), llama.cpp support, and GGUF model import — and more

[!IMPORTANT]

Disclaimer

Box began as a fork of Google AI Edge Gallery and is not affiliated with or endorsed by Google LLC. Google branding has been replaced throughout. Box has since diverged substantially from upstream — active merging with upstream stopped some time ago, and upstream has itself since adopted features that originated in Box. Box now carries roughly 50+ features not present in upstream Google AI Edge Gallery. Credit for the original underlying platform goes to Google and the original contributors.

Changelog v1.0.7 – v3.4.5

VersionFeatureDetails
v3.4.5Google Firebase removedFirebase Analytics and Firebase Cloud Messaging came in with the original Google AI Edge Gallery fork and were never used by Box. There was no configuration file for them, so they could not start up, collect anything or send anything — and analytics was switched off in the manifest on top of that. They are now gone from the app entirely rather than merely disabled, along with 20 dormant tracking calls and six Google entries in the app manifest. Nothing you can see or do in Box changes.
v3.4.5Google usage logging switched offRemoving Firebase surfaced a second piece of Google code: Clearcut, a usage-logging transport that arrives inside ML Kit (used for background removal, face detection and reading text from images), so it did not leave with Firebase. It sends over an ordinary HTTPS connection rather than through Google Play services — meaning a de-Googled phone does not stop it — and ML Kit provides no setting to turn it off. Box now disables it at the source: with nothing registered to receive them, events are discarded before they are even written down. The ML Kit features themselves are unaffected.
v3.3.5Pixel 11 / Tensor G6 accelerationTwo models rebuilt for the Tensor G6 in the Pixel 11 — Gemma-4-E2B-it (Tensor G6) (3.3 GB, 32K context, text/image/audio) and Gemma 3 1B-IT (Tensor G6) (2.0 GB, text only). Both run on the phone's dedicated AI hardware rather than the GPU, and are substantially faster there than the standard build. They appear automatically on a Pixel 11 and are hidden everywhere else. The Pixel 10 / Tensor G5 path is unchanged.
v3.3.5Themes for colour blindnessThree new Ptyxis+ palettes — Deuteranopia, Protanopia and Tritanopia, one per common type of colour blindness. Each palette's colours are drawn from the standard published colour-blind-safe sets (Okabe–Ito, Paul Tol, IBM) and were chosen by simulating that specific condition and maximising the separation between the palette's worst-matched pair, so no two colours collapse into one. An ordinary theme put through the same check usually has at least one indistinguishable pair — blue and cyan being the classic. Plus Seafoam Pastel, a new regular palette, bringing Ptyxis+ to 42.
v3.3.5New font — AliceA warm, readable serif, selectable under Settings → Font. Egyptian Hieroglyphs has been removed; if you had it selected Box falls back to the default font on its own.
v3.3.4Restore — deblur and denoise your photosA new tile under Image. Deblur sharpens shots ruined by camera shake or a moving subject; Denoise cleans the speckled grain phones produce in dim light without smearing detail away. Both run on the GPU via LiteRT and are bundled in the app, so there is nothing to download and it works with no connection. Large photos are processed in overlapping tiles and stitched, so the result comes back at the size you put in.
v3.3.4Flatten — straighten a photo of a pagePhotograph a receipt, a book page or a form at an angle and Box will flatten the curl out so the text sits straight. Downloaded on first use (180 MB) rather than bundled, since it is a specialist tool.
v3.3.4Portrait SketchTurn a portrait photo into a pencil line drawing. Works best on a single, well-lit face looking at the camera. Downloaded on first use (168 MB).
v3.3.4App language — Deutsch & 简体中文German and Simplified Chinese join French, Portuguese and Português (Brasil) in Settings → Language. Both were translated properly rather than merged from upstream, so every screen that can follow your language setting now does. A good deal of Box's text is still written directly into the app rather than being translatable, though, so English still shows in places — in every language.
v3.3.4Five light terminal themesThe Ptyxis theme picker was dark-only. It now also offers Belafonte Day, Everforest Light, GitHub Light, Solarized Light and Xterm Light, taken from the same upstream palettes as the existing 33. Every existing theme is unchanged.
v3.3.4Model list fixesThe model switcher inside a chat could not be scrolled, so with 41 models everything past the fold was unreachable — fixed. Models you have already downloaded now sort to the top of the list, in the switcher and in the Models browser, so you are no longer scrolling to the same place every time.
v3.3.4Gemma 4 12B output fixGemma 4 12B spliced stray <image|> markers into ordinary replies. Its bundle carries a multimodal checkpoint whose image support the runtime has not enabled yet, and the runtime rendered those unused media tokens as visible text. They are now filtered out. Applies only to the two models that can produce them; every other model's output is untouched.
v3.3.4Advanced Protection Mode supportIf you have Android's device-wide Advanced Protection switched on, Box tightens up to match: the biometric lock is forced on and cannot be switched off, MCP is forced off and cannot be switched on, and the tamper check terminates outright instead of showing a dismissible screen. Nothing changes if you do not use Advanced Protection. Network access is deliberately left alone so model downloads still work.
v3.3.4Two models removedQwen 3.5 0.8B and Polaris 4B Preview are gone — neither could actually run in Box. Qwen 3.5 needs a newer LiteRT-LM than Box ships (its architecture is not supported by the current runtime), and Polaris asks for roughly 12 GB of GPU memory. Both were failing rather than merely slow.
v3.3.4Under the hoodDeclares android.hardware.npu, which Android 17 requires of apps that use the NPU.
v3.3.3🎨 Bonsai Image 4B — the new recommended image generatorA ternary-weight build of FLUX.2 [klein] running fully on-device via LiteRT. 512×512 output — double klein's tile — from a smaller download: ~4.3 GB against klein's 7.4 GB and Z-Image's 10.6 GB. Four steps, guidance-free, no internet at any point. It runs on the CPU (its 2.27 GB diffusion graph is too large for the GPU delegate), so allow a couple of minutes per image and around 8 GB of RAM. Send the result through Upscale → EDSR ×4 for a 2048×2048 image.
v3.3.3⚡ Faster GGUF chat and faster image generationllama.cpp, whisper.cpp, stable-diffusion.cpp and ggml all updated to current builds — months of upstream work in one go. GGUF chat and Stable Diffusion image generation are both noticeably quicker on the same phone with the same model, and there is nothing to configure. Whisper transcription and Gemma/LiteRT chat behave exactly as before.
v3.3.3🌍 App language — French & PortugueseNew Settings → Language picker: System, English, Français, Português and Português (Brasil). Box uses Android's per-app language support, so your choice is remembered by the system. Coverage is partial for now (following upstream) — translated screens follow your selection, the rest stays in English.
v3.3.315 new modelsFour new vision models for Ask Image: SmolVLM2-2.2B (1.5 GB), SmolVLM2-500M (just 0.36 GB), InternVL3.5-2B and InternVL3-2B — the InternVL pair are strong at reading text in photos. New chat and reasoning models: Qwen3.5-0.8B (hybrid attention, so memory stays flat as the conversation grows), Phi-4-mini-reasoning, Polaris-4B Preview, Nanbeige 4.2 3B, SmolLM3-3B, Jan-nano, Ministral 3 3B in both Instruct and Reasoning builds, OLMo-2-1B Instruct (fully open weights, data and training code), Granite-4.0-H-1B and LFM2.5-1.2B-JP for Japanese. Gemma 4 12B is now a 560 MB smaller download for exactly the same capability.
v3.3.3✍️ Model descriptions rewritten in plain EnglishAround 70 cards on the download page rewritten. Each now opens with what the model actually does and when to pick it, with the technical specifications kept at the end for those who want them. Chip-specific builds say "For Pixel 10 only" or "For Snapdragon 8 Elite phones only" up front, so it is obvious which download suits your phone. Licences, RAM warnings and Gemma Terms of Use notices are all preserved.
v3.3.3FixesChat now works properly on de-Googled Android (custom-rom-support build): on GrapheneOS, AOSP, crDroid, LineageOS and similar, AI Chat silently fell back to the CPU whichever accelerator you picked, and the Tensor G5 model would not load at all — failing with "Input tensor not found" — even though the Benchmark screen ran the very same model on the TPU. Both are fixed. Large downloads now resume by themselves instead of staying stuck until you closed and reopened Box.
v3.3.3Under the hoodAGP 9.3.1, Gradle 9.6.1, Kotlin 2.3.10 and 28 library updates. Debug logging is stripped from release builds and native debug symbols are no longer packaged. LiteRT and LiteRT-LM are deliberately held at their current versions.
v3.3.2Downloads fixedModel downloads are reliable again after 3.3.1 — no more failing mid-download or stalling at 100%. A previously stuck model downloads normally on the first try.
v3.3.2GGUF GPU crash fix (really this time)The Snapdragon GPU crash fix from 3.3.1 now actually ships in the build.
v3.3.2Biometric lock + database encryptionThe biometric app lock works alongside database encryption again — the two are independent, and the app re-locks when reopened.
v3.3.1Live Translator (NEW, Sound tab)Two people, two languages — tap your button, speak, and the other person reads and hears it in their language. Runs on your installed Gemma audio model (E2B/E4B), each phrase translated on its own for flat latency. 24 languages, fully offline.
v3.3.14 new modelsGranite 4.0 350M (IBM's tiny fast tier, 468 MB), MiniCPM5-1B in int8 and int4 builds, and experimental Gemma 4 26B (A4B) — Google's mixture-of-experts Gemma for 16 GB+ RAM devices.
v3.3.1FixesGGUF models no longer crash on GPU on some Snapdragon devices (Adreno driver quirk). Rotating or folding the phone no longer unloads the model. Custom-ROM: TPU/GPU chat works again on de-Googled devices (GrapheneOS).
v3.3.0🦯 Box Assist — a camera that talks (NEW)Spoken camera assistance for blind and low-vision users, under the Core tab. Live mode calls out people, obstacles and objects with how close they are; Reading mode reads mail, labels and menus aloud; Describe mode describes the scene, spoken as it thinks; voice questions — double-tap, ask out loud, and Box answers against what the camera sees. One download bundles everything (vision models + the Describe brain + speech recognition). Continuous autofocus with pre-capture focus sweeps, automatic flashlight in the dark, a blur check on Reading, physical volume-button controls, hold-to-repeat, TalkBack coexistence, screen never times out, and a launcher long-press shortcut straight into it. Fully offline.
v3.3.0⚡ GGUF engine rebuilt — real GPU accelerationThe llama.cpp engine got a ground-up overhaul: full Vulkan GPU offload via the CPU/GPU chip in any GGUF chat, a massively faster CPU mode (a flaw routed CPU prompt processing through the GPU — 0.7 → 21 tok/s on a Pixel 6a), instant replies (weights read up front, reopened chats replay their history during the loading screen), a tokens/sec stat under every GGUF reply, a new Settings → GGUF Models panel (context size, CPU threads, GPU layers, mmap, mlock, Q8 KV cache), sturdier imports with byte-verification, and automatic GPU→CPU retry. llama.cpp updated to a current build.
v3.3.0🎨 On-device image generation — FLUX.2 klein & Z-Image TurboTwo full text-to-image diffusion models running 100% on-device via LiteRT: FLUX.2 klein (4B) — photorealistic images in 4 steps (~7.4 GB download) — and Z-Image Turbo (9 steps), which shares nearly a gigabyte of files with klein so Box doesn't download them twice. Multi-gigabyte downloads now resume without refetching finished files, progress bars show honest totals, and a model only shows "downloaded" when every file is actually present.
v3.3.0🔍 Five new vision models — bundled, work instantlyIdentify now hosts four model families in one picker: MobileNet V2, MobileNet V3 Large (with a Pixel Tensor G5 NPU variant), PlantNet (identify 1,081 plant species from a photo) and DM-Count crowd counting. New Erase tile — paint over anything in a photo and MI-GAN inpainting removes it (brush size, iterative erase, save to gallery). Upscale gains EDSR ×4. All bundled in the APK — no download, fully offline.
v3.3.0📱 Android 14 supportMinimum Android version lowered from 15 to Android 14 — Box now installs on a whole generation more of phones.
v3.3.0Fixes & polishBox Assist: fixed a first-open black screen (camera and mic permission requests raced each other) and made repeat scene descriptions as fast as the first. Fixed a case where an already-loaded model would never signal "ready", leaving features waiting forever. Download cards show accurate total sizes before you tap.
v3.2.0🎵 On-device music & sound generationMake music and sound effects from a text description — completely offline, nothing leaves your phone. Three tiers under the new Sound tab: SoundGen (quick clips & sound effects in seconds), SoundGen HD (higher-quality audio up to ~24s), and SoundGen HD Long (full pieces up to ~3 minutes). Set the length, then play, save, or share the result. The generator for each tier downloads on first use, then runs entirely on-device.
v3.2.0Identify — on-device image recognitionPoint Box at a photo and it tells you what's in it — 1000+ everyday objects, animals and scenes. Pick from your gallery or take a new shot. Fully offline, hardware-accelerated on supported devices.
v3.2.0Tabs reorganised — Sound & CoreClearer home tabs: Sound groups the audio features, Core groups chat & assistant.
v3.2.0Chat remembers on reopenReopening a conversation now replays recent context to the model, so it picks up where you left off — new chats still start fresh.
v3.1.0NPU now works on Snapdragon & MediaTek — for the first timeThis is the first Box build where on-device NPU acceleration actually runs on Snapdragon and MediaTek phones. Previous builds shipped the NPU models but crashed on load. Box now ships the Qualcomm and MediaTek NPU dispatch libraries rebuilt to match the LiteRT runtime plus an updated Qualcomm AI stack (QNN 2.47), with per-vendor builds so each phone loads the correct driver — NPU chat and benchmarking now run on those devices. The Pixel / Tensor G5 path is unchanged. (#81, #83, #88)
v3.1.0Smoother NPU chat on small modelsLong conversations on the Gemma 3 1B NPU model no longer abruptly stop or error when the context fills — Box slides the context window so the chat keeps going. Added safeguards so the small NPU model doesn't get stuck repeating itself or return empty replies. (Snapdragon / MediaTek NPU only — Tensor G5 and GPU/CPU are untouched.)
v3.1.0Fix — NPU benchmark crash (#81)Benchmarking an NPU model no longer crashes.
v3.1.0PolishNew animated "Initializing model" loading screen; removed the "Experimental" tag from Mobile Actions; tidied up model descriptions.
v3.0.0Major UI overhaul — Material 3 ExpressiveA top-to-bottom interface refresh. The app now moves with spring-physics motion: home cards bounce in and respond to taps, chat messages rise and fade in as they arrive, and screen transitions use Material 3 slide-and-fade. The jump to 3.0.0 reflects how much of the UI changed — the models and engines are unchanged.
v3.0.011 new themesA set of terminal-inspired palettes — Fairy Floss, Nord, Bim, Borland, C64, Cobalt Neon, Grass, Homebrew Ocean, Mono Amber, Mono Red, and Synthwave — selectable from a new dropdown in Settings, alongside the existing System, Light, Catppuccin, and Dracula themes.
v3.0.0Custom app & chat fontsChoose from 13 bundled font families (Nunito plus Cormorant Garamond, DotGothic16, IBM Plex Mono / Serif, Instrument Serif, Playfair Display, Press Start 2P, Quicksand, Space Grotesk, Turret Road, Viaoda Libre, and more), each previewed in its own typeface — with an optional separate font just for chat messages.
v3.0.0Text-size sliderScale text across the whole app and chat from 0.8× to 1.4×, on top of your system font size.
v3.0.0Themed app iconWith "Themed icons" enabled in your launcher, the Box icon now tints to your system Material You colours.
v3.0.0Settings, reorganisedThe long settings list is now grouped into smooth, collapsible categories — Appearance, Privacy & Security, Network & Tools, Chat & Voice, and About.
v3.0.0Theme-aware task screensOpen any task (Chat, Diffusion, Voice…) and the background now follows your selected theme with the same accent tint as the home screen — no more flat black behind a colourful theme. Cleaner, icon-free task headers, a tidy box-shaped menu button, and a new Material 3 wavy download-progress indicator.
v3.0.0Fix — NPU crash on Snapdragon & MediaTek (#82, #83)NPU models could hard-crash on load on non-Pixel devices (e.g. Galaxy S26 Ultra, Xiaomi 14T Pro) because the wrong hardware dispatch library was being loaded. Box now selects the correct Qualcomm / MediaTek runtime per device. The Pixel 10 / Tensor G5 path is unchanged and re-verified.
v3.0.0Fix — Settings flash & jankChanging the text-size slider no longer flashes the home screen behind Settings, and opening Settings or expanding a category no longer jumps — the dialog is now fixed-size and animates its contents internally.
v2.0.2New model tier — Gemma 3 270MBrand-new ultra-lightweight model (~460–555 MB) — fast and low-RAM, ideal for quick tasks on modest devices. Ships dedicated NPU builds for Snapdragon (SM8550 / 8650 / 8750 / 8750-AB / 8850) and MediaTek Dimensity (MT6991 / MT6993).
v2.0.2New models — Gemma 3 1B-IT with broad NPU coverageGemma 3 1B now ships dedicated on-device NPU builds across Snapdragon (SM8550 → SM8850, incl. the Samsung SM8750-AB) and MediaTek Dimensity (MT6989 / 6991 / 6993), plus a universal GPU/CPU build. Each device automatically downloads the build that matches its chip.
v2.0.2New models — Gemma 3n E2B & E4B (multimodal)Text, image and audio input, up to 32K context, with Gemma 3n's selective-parameter architecture. Run on GPU/CPU on every device; NPU-accelerated on MediaTek (MT6993).
v2.0.2Samsung Galaxy S26 Ultra (SM8850) NPU modelsAdded SM8850 ("Snapdragon 8 Elite Gen 5") allowlist keys across the new Gemma 3 1B and 270M entries, so dedicated NPU models now appear and run on the S26 Ultra.
v2.0.1Fix — Snapdragon 8 Elite NPU crash (SM8750 / SM8750-AB)The audio sub-graph was incorrectly routed to the NPU on all SM8750 devices, causing an instant hard crash (SIGABRT) when loading the Snapdragon NPU model — no error popup, just an immediate exit. Audio always uses CPU regardless of the primary backend, matching upstream behaviour. Fixes Red Magic NX799J, iQOO 13, and any other SM8750 or SM8750-AB device.
v2.0.1Fix — Samsung Galaxy S25 / S26 Ultra NPU models not listedSamsung's "Snapdragon 8 Elite for Galaxy" variant reports SM8750-AB as its SoC identifier, not SM8750. The model allowlist only matched sm8750, so dedicated NPU models were invisible to all S25 and S26 Ultra users.
v2.0.1New model — Gemma 4 E2B (Qualcomm QCS8275 / Dragonwing IQ8)Added an NPU model entry for the Qualcomm QCS8275 SoC. Appears automatically on matching hardware.
v2.0.0Google Tensor G5 (Pixel 10) accelerationGemma now runs on the Pixel 10's Tensor G5 TPU, not just the GPU. Supported models route to the TPU automatically and expose a dedicated TPU option in the accelerator picker.
v2.0.0MediaTek NPU supportBundled the MediaTek dispatch runtime and added the first models that run on MediaTek Dimensity neural engines.
v2.0.0New modelsGemma 4 E2B (Tensor G5) and Gemma 4 12B (GPU); Gemma 3 1B-IT (Tensor G5), Gemma 3n E2B (MediaTek, multimodal) and Qwen3 0.6B (MediaTek)
v2.0.0Face Recognition — on-device & encryptedNew tool in the image section: detect, enroll and name people, then recognise them in photos or live from the camera, fully offline. Multi-sample enrollment with face alignment, capture-to-add, an on-screen face mesh, and a settings panel (match strictness, front camera, show %, clear all). All face data is encrypted on-device (SQLCipher) and never leaves the phone — opt-in and user-enrolled only.
v2.0.0New Light theme + theme-aware homeA crisp, wallpaper-independent Light theme, and the home background now follows your selected theme (System / Light / Catppuccin / Dracula) instead of always being black.
v2.0.0Gemini Nano Hub on custom-ROMThe full Gemini Nano hub (Summarize / Proofread / Rewrite / Describe / Chat / Speech) is now included in the custom-rom-support build too, degrading gracefully on devices without AICore (ML-Kit vision tools still work).
v2.0.0Nano document-attach crash + leak fixesFixed a crash when attaching a document in Summarize/Proofread/Rewrite (the file picker could be hijacked by the photo picker on Android 14+) — now uses the proper document picker with a clean fallback. Also fixed GenAI service/memory leaks when switching between Nano features.
v2.0.0Copy button on code blocksFenced code blocks in chat now render with a language label and a one-tap Copy code button.
v2.0.0SenseVoice in ChatThe chat mic now works with a loaded SenseVoice model (priority Whisper → SenseVoice → system) instead of dead-ending when no Whisper model is present.
v2.0.0Speculative decoding in chatSpeculative / Multi-Token-Prediction decoding is available for Gemma 4 in chat (off by default).
v2.0.0Fix #69 — agent mode with text-only modelsAgent mode no longer force-loads vision on models that don't support it, which previously blocked text-only imported models entirely.
v2.0.0Fix #67 — correct installed versionAligned versionName with the public version, so Obtainium / Android's "App version" report the right number (no more false "update available"). This is why the release jumps to 2.0.0.
v2.0.0Smaller downloadNative libraries are now compressed inside the APK — the main build drops from 400 MB+ to ~278 MB (they're extracted on install).
v1.0.12SenseVoice — multilingual speech-to-textNew card in the Voice tab. Transcribes Chinese, English, Japanese, Korean and Cantonese fully offline, roughly 5× faster than Whisper on CPU. Live "listening" preview while you talk, a multi-message transcript log (copy / delete / clear), language picker, punctuation & number formatting, and optional emotion / audio-event tags. (#68)
v1.0.12Supertonic — multilingual text-to-speechNew card in the Voice tab. Lightweight (~66M param) on-device speech synthesis in English, Korean, Spanish, Portuguese and French, with multiple built-in voices and adjustable speed. Fully offline — text never leaves the device.
v1.0.12AI Image Upscaling (super-resolution)New Upscale tool in the image tab. Enhance and enlarge any photo 4× on-device and save it to your gallery. Three models bundled in the app — XLSR (fast), Real-ESRGAN General (balanced), Real-ESRGAN x4plus (quality) — run via LiteRT, no download required. Photos are auto-rotated (EXIF-aware) before upscaling.
v1.0.12Gemini Nano Vision — visual overlays (main)Pose detection now draws a skeleton overlay and Face Mesh a 468-point mesh directly on the camera preview and still images (previously text-only). Added copy buttons on every vision result, an adjustable live refresh rate (Fast / Balanced / Slow / Power-saver) with a Freeze/Resume toggle, front/rear camera switching on all modes, and image upload from your gallery.
v1.0.12Models browser organised by typeThe model list is now grouped into Language models / Speech-to-Text / Text-to-Speech / Image generation / Other instead of one flat alphabetical list.
v1.0.12New language modelsAdded TinyLlama 1.1B, Phi-4-mini, TinySwallow 1.5B, VibeThinker 1.5B, and Qwen3 8B to the download list.
v1.0.12Markdown & LaTeX rendering overhaul (#42)Headers, bullet/numbered lists and bold text now render correctly even when mixed with inline math on the same line; bold that spans a math expression no longer shows literal **; wide display equations scroll instead of being clipped.
v1.0.12Clearer model guidance + UI cleanupGemma 4 E2B labelled "Recommended", E4B "Best overall for flagship devices," with cleaned-up model descriptions. Removed promotional banners/links from the MCP and Agent screens (sample-prompt chips kept).
v1.0.12Fix #59 — Snapdragon NPU crashVision/audio sub-backends now follow the primary backend on the NPU path, fixing hard crashes on some Snapdragon devices.
v1.0.12Fix #61 — leftover model filesOrphaned model-version directories are cleaned up after app updates.
v1.0.12Fix #65 — GrapheneOS speech hangRestored the SpeechRecognizer availability gate (custom-rom-support build).
v1.0.12Fix — config dialog crashOpening the model settings dialog on small-context-window (<2000) models no longer crashes.
v1.0.12Android SDK 37 + deeplink fixUpdated compile/target SDK to 37 and fixed the notification tap deep link.
v1.0.11MCP server supportThe Agent tab can now connect to external Model Context Protocol servers (e.g. gitmcp.io/<owner>/<repo>) and give the model access to remote tools. Off by default — enable in Settings, add a server URL, accept the disclaimer. Every tool call fires a per-call permission dialog (Allow once / Always allow / Deny). Hard Offline Mode disables MCP.
v1.0.11"Agent Skills" renamed to "Agent"Reflects the addition of MCP tools alongside the existing 20 built-in skills. Internal IDs unchanged.
v1.0.11Broader NPU init crash recovery (main)Snapdragon 8 Elite / Vivo OriginOS users (e.g. iQOO 13) reporting hard crashes on NPU model open now fall back silently to GPU instead. Any catchable NPU init exception is recovered, not just TF_LITE_AUX.
v1.0.11Pixel 8/9 TPU labelTensor G3 / G4 devices now show the TPU accelerator label alongside Pixel 10 (isPixelDevice() broadened from isPixel10()).
v1.0.11Smoother streaming renderBufferedFadingMarkdownText two-layer crossfade reduces markdown re-render jank during token streaming.
v1.0.11Chat scroll performancesnapshotFlow + derivedStateOf translated to Box's LazyColumn. Significantly fewer Compose recompositions per generated token.
v1.0.11ChatGPT-style chat layoutUser and assistant messages both left-aligned, restoring Box's original look.
v1.0.11Downloaded-model tick iconOnce a model is on device, the model picker chip and Model Manager show a filled-circle tick instead of the download-arrow icon.
v1.0.11Gemma 4 model hashes refreshedGemma 4 E2B / E4B / E2B-Snapdragon entries updated to upstream's latest commits (6e5c4f1e… / 28299f30…).
v1.0.11R8 keep rule for tool callsRelease builds preserve @Tool method names on every ToolSet subclass — MCP and Agent skills now work in release APKs (was silently broken).
v1.0.11Upstream merged to 1.0.15Internal versionName bumped to match upstream gallery 1.0.15 (cherry-picked over multiple sessions; chat history, model schema, and other heavily-customised Box paths preserved).
v1.0.10Gemini Nano hub6 on-device ML Kit features powered by Gemini Nano on Pixel 9+ (via AICore, NPU/TPU-accelerated): Summarize, Proofread, Rewrite, Chat, Describe Image, and Speech-to-Text. First use triggers an automatic background download of Gemini Nano (~1–2 GB via AICore).
v1.0.10Nano Chat — multi-sessionPersistent multi-turn chat with Gemini Nano. Sessions are stored in the existing encrypted SQLCipher database, auto-titled from the first message, and fully resumable. Sessions can be renamed or deleted. Long-press any bubble to copy.
v1.0.10Document attachment in NanoProofread and Rewrite now accept attached documents (PDF, TXT, MD) — content is read and passed to Gemini Nano as context.
v1.0.10Live camera in Describe ImageGallery tab + Live Camera tab. Camera tab binds an ImageCapture use case — tap Capture to send the current frame to Nano for description.
v1.0.10Background RemovalNew tool powered by ML Kit Subject Segmentation (main branch). One tap removes the background from any photo with a transparency-preserving PNG output. Includes a "Trim transparent edges" toggle. Save or share the result.
v1.0.10Catppuccin + Dracula themesThree-way theme picker in Settings: System (Material You) / Catppuccin (14 accents) / Dracula (7 accents). Accent colour persists across restarts with no first-frame flicker.
v1.0.10Tap jacking protection toggleNew toggle in Settings (on by default) — filterTouchesWhenObscured blocks touch events when an overlay is detected, preventing tap-jacking attacks.
v1.0.10Accessibility data sensitivity toggleNew Settings toggle hides app content from untrusted accessibility services. Off by default (note: incompatible with TalkBack).
v1.0.10LaTeX in table cellsInline math inside markdown table cells no longer wraps across multiple lines. Uses Compose InlineTextContent to embed math as a single placeholder inside Text().
v1.0.10Import button simplifiedHome screen import button label shortened to just "Import" (removed "GGUF · LiteRT" subtitle).
v1.0.10NPE crash fixFixed a null-pointer crash on startup and on Retry caused by a broken fallback comparator in groupTasksByCategory.
v1.0.9Document Q&ANew RAG pipeline: import PDFs and ask questions grounded in the document. Uses MiniLM embeddings (on-device, LiteRT) for chunk retrieval — model only sees the relevant passages. Every answer cites the source chunks it used.
v1.0.9Model picker in Document Q&AChoose which downloaded LLM handles answering — defaults to first available, switchable mid-session.
v1.0.9Kokoro TTS (English)Single Kokoro model (csukuangfj/kokoro-en-v0_19, ~346 MB) replaces broken individual-voice entries. Correct tensor shapes and metadata — works first time.
v1.0.913 Piper voices8 new voices: LibriTTS-R, HFC Female, HFC Male, Arctic (US English); Thorsten (German); UPMC (French); MLS 10246 (Spanish); Huayan (Chinese Mandarin). 13 total across both branches.
v1.0.910 Whisper modelsExpanded from 3 hardcoded to 10: Tiny, Base, Small, Medium, Large-v3-Turbo, and Large-v3 — each in multilingual and English-only variants. Shared across Audio Scribe and Voice Input.
v1.0.9Gemma-4-E2B-it (Snapdragon 8 Elite)NPU-optimised variant added to the model allowlist — visible only on SM8750 devices.
v1.0.9Fix #46 — Audio Scribe OOM crashReplaced boxed List<Float> (~16 bytes/sample) with a primitive growing FloatArray (4 bytes/sample). 30-min audio at 16 kHz no longer causes ~460 MB excess allocation.
v1.0.9Fix #47 — TTS silent with non-Amy voiceAuto-init and GrapheneOS TTS fallback now filter by download status before selecting a voice model (custom-rom-support only).
v1.0.8Saved System PromptsSave, name, and reuse system prompts from the model settings dialog. Tap to apply, swipe to delete.
v1.0.8Restore DefaultsNew button in model settings resets all sliders (temperature, top-K, top-P, max tokens) back to defaults in one tap.
v1.0.8System prompt actually appliedChanging the system prompt mid-session now correctly resets the conversation with the new instruction — previously saved in UI but not passed to the model.
v1.0.8Markdown fix in math responsesPlain-text segments in chat bubbles now render through the Markdown pipeline, fixing broken formatting in responses that mix text and LaTeX math.
v1.0.8Randomised inference seedEach conversation now uses a unique random seed for more varied outputs on CPU backend.
v1.0.8GPU determinism root cause foundLiteRT LM v0.11.0 hard-caps max_top_k: 1 on devices without a GPU sampler, forcing greedy decoding. Switch to CPU for varied outputs. Reported upstream as issue #817.
v1.0.7Gemma 4 E2B & E4B updatedModel files refreshed on HuggingFace — new commit hashes, smaller sizes, same multimodal capabilities.
v1.0.7Speculative decoding / MTPMulti-Token Prediction reads capability from the model file itself. Gemma 4 E2B reaches 66–91 tok/s on Galaxy S26 Ultra (GPU + spec) vs 52 tok/s plain GPU.
v1.0.7Sustained Performance ModesetSustainedPerformanceMode(true) locks clocks during inference — no mid-conversation thermal throttling on long generations.
v1.0.7Benchmark spec decoding toggleBenchmark screen shows a speculative decoding toggle for supported models.
v1.0.7AI Chat app shortcutLong-press the Box icon → AI Chat jumps straight into chat, even from a cold start.
v1.0.7In-app update checkerSettings → Check for updates — fetches the latest GitHub release and offers a direct download link for your variant.
v1.0.7Model import from listWhisper and TTS models can now be imported directly from the model list.


Related

Built OfflineLLM first — a privacy-first Android chat app with a pure llama.cpp backend.


What is Box?

Box Header

Box is an Android app for running AI entirely on-device — chat, voice mode, image generation, image upscaling, speech-to-text, text-to-speech, document analysis, and vision, all without a network connection. It inherits the full feature set of the upstream Google AI Edge Gallery and layers on top: encrypted conversations, biometric lock, hard offline mode, and three additional native inference engines (llama.cpp, stable-diffusion.cpp, whisper.cpp) alongside LiteRT.

Box: On-Device AI. No Cloud. No Compromise.

What makes Box unique? You can sit at your desk, tap two buttons, and have a real flowing voice conversation with an AI — no wake word, no account, no server, no subscription. It listens, thinks, and speaks back sentence by sentence before it's even finished generating. Point the camera at something and ask about it out loud. The AI sees it and answers. All of it runs on the phone in your hand, completely offline, faster than you'd expect.


Screenshots

Box
Home — Chat
Box
Home — Diffusion
Box
Home — Voice
Box
AI Chat
Box
Model Config
Box
Model Manager
Box
Text to Speech
Box
Voice Input
Box
Whisper Scribe
Box
Image Generation
Box
Gemini Nano Hub
Box
MCP — Add Server
Box
Settings — Theme & Security
Box
Settings — Behaviour & MCP
Box
Settings — About

[!NOTE]

What Box adds on top of upstream

Box started off as a fork of Google AI Edge Gallery. The upstream project is excellent — Box layers on additional capabilities and features not present in upstream.

AreaWhat Box adds
Inference enginesllama.cpp (GGUF LLMs, full Vulkan GPU offload), stable-diffusion.cpp (image gen), whisper.cpp (STT) alongside LiteRT
Model importImport any local GGUF file — not limited to the curated download list
NPU / TPUAll Snapdragon / Tensor / MediaTek variants bundled in one APK (upstream ships per-SoC)
Box AssistSpoken camera assistance for blind and low-vision users — Live object/proximity callouts, Reading (OCR aloud), Describe (scene answers, spoken as generated), voice questions. One bundled download, autofocus + auto-flashlight, volume-button controls, TalkBack-friendly, fully offline
Voice mode / Vision modeFree talk (continuous hands-free loop) and Vision talk (live camera + voice)
Image generationOn-device Stable Diffusion via GGUF, plus Bonsai Image 4B (512×512, recommended), FLUX.2 klein (4B) and Z-Image Turbo diffusion via LiteRT
Image recognitionIdentify: MobileNet V2 / V3 Large (+ Tensor G5 NPU variant), PlantNet (1,081 plant species), DM-Count crowd counting — bundled, offline
Erase (inpainting)Paint over anything in a photo and MI-GAN removes it — brush size, iterative erase, save to gallery (bundled, offline)
Music & sound generationGenerate music and sound effects from a text prompt, fully offline — quick clips, higher-quality audio, or long-form pieces up to ~3 minutes (Sound tab)
Image upscalingAI super-resolution — enlarge any photo 4× on-device (XLSR / Real-ESRGAN / EDSR via LiteRT), models bundled, fully offline
Speech-to-textOn-device Whisper STT, plus SenseVoice for fast multilingual transcription (Chinese / English / Japanese / Korean / Cantonese, ~5× faster than Whisper)
Text-to-speechSupertonic multilingual on-device TTS (5 languages, multiple voices) alongside Piper / Kokoro
Document analysisAttach text files (.txt, .md, .csv, .kt, etc.) directly in chat
Document Q&ARAG pipeline: import PDFs, embed with MiniLM on-device, ask questions grounded in document content — answers cite their source passages
Gemini Nano6 on-device ML Kit features (Summarize, Proofread, Rewrite, Chat, Describe, Speech) — entirely on-device via AICore on Pixel 9+/10 and recent Samsung / Xiaomi / OnePlus / OPPO / vivo flagships (both branches as of v2.0.0). Vision modes add live camera + still-image analysis with visual overlays (pose skeleton, 468-point face mesh)
Face RecognitionOn-device, encrypted face recognition (both branches) — enroll and name people, then recognise them in photos or live from the camera. Multi-sample enrollment with alignment, capture-to-add, face-mesh overlay, SQLCipher-encrypted storage, fully offline and opt-in
Background RemovalML Kit Subject Segmentation — remove backgrounds from photos, output a transparency-preserving PNG (main branch)
Chat historyPersisted to a SQLCipher-encrypted Room database, resumable across sessions
SecurityBiometric app lock, hard offline mode, prompt sanitisation, audit log, tap jacking protection, accessibility data sensitivity
ThemesCatppuccin (14 accents), Dracula (7 accents), a bright Light theme, and Material You — picker in Settings, with the home screen tinted to match the active theme
Agent (skills + MCP)20 built-in skills (upstream has 9) plus Model Context Protocol — connect to remote MCP servers and give the model real tools, with per-call permission prompts
Math renderingLaTeX expressions rendered as Unicode in chat, including inside markdown table cells
App shortcutsLong-press icon → AI Chat or Box Assist for instant cold-start navigation
In-app updatesSettings → Check for updates — compares against latest GitHub release, downloads correct variant

Core Features

Local Chat

Multi-turn conversations with on-device LLMs. Import any GGUF model or download LiteRT models from the built-in list. Supports Thinking Mode on compatible models. Full markdown rendering with LaTeX math support — Greek letters, operators, fractions, and notation are rendered as Unicode symbols. Conversations are persisted and resumable.

Recommended models: We highly recommend Gemma 4 E2B or Gemma 4 E4B (LiteRT) as your primary models — best-tested, support vision, voice, and documents, and run efficiently with GPU/NPU acceleration. Available to download directly in the app.

With Gemma 4 E2B / E4B selected, the chat input expands to a full multimodal interface:

  • 📎 Attach documents (.txt, .md, .csv, .json, .py, .kt, and more) — content is injected into context automatically
  • 🎙 Record an audio clip or pick a WAV file to speak your question
  • 📷 Take a photo or pick from album for visual Q&A

Box Assist — Spoken Camera Assistance

Built for blind and low-vision users, and useful to anyone who wants a talking camera. Live mode calls out people, obstacles and objects around you with how close they are; Reading mode reads mail, labels, menus and signs aloud; Describe mode answers "what's in front of me?" in a couple of spoken sentences — streamed aloud as the model generates; double-tap and ask anything out loud and Box answers against what the camera sees. One download includes everything (vision models, the Describe brain, and speech recognition). Continuous autofocus with a focus sweep before every capture, automatic flashlight when it's dark (Box tells you), a blur check so Reading waits for a sharp frame, physical volume-button controls, hold-to-repeat, TalkBack coexistence, and a launcher shortcut straight into it. Everything runs on-device.

Local Diffusion

On-device image generation powered by stable-diffusion.cpp. Runs Stable Diffusion 1.5 in GGUF format fully offline — no API key, no cloud. Configurable steps, CFG scale, seed, and image size presets. Save generated images directly to your gallery. Import your own GGUF diffusion models.

Image Generation — Bonsai, FLUX.2 klein & Z-Image Turbo

Three full text-to-image diffusion models running 100% on-device via LiteRT.

Bonsai Image 4B — recommended. A ternary-weight build of FLUX.2 [klein] that produces 512×512 images in 4 steps from a ~4.3 GB download, guidance-free. It runs on the CPU rather than the GPU — its 2.27 GB diffusion graph is too large for the GPU delegate — so allow a couple of minutes per image and around 8 GB of RAM. Send the output through Upscale → EDSR ×4 for a 2048×2048 image.

FLUX.2 klein (4B) generates photorealistic images in just 4 steps (~7.4 GB download); Z-Image Turbo runs in 9 steps (~10.6 GB) and shares nearly a gigabyte of its files with klein, so Box is smart enough not to download those twice. Interrupted multi-gigabyte downloads resume without refetching finished files.

All three produce noticeably better output than the bundled Stable Diffusion GGUF models.

Music & Sound Generation

…view the full README on GitHub.

// faq

What is Box?

The most advanced, fully offline client-side AI suite on Android today.. It is open-source on GitHub.

Is Box free to use?

Box is open-source under the NOASSERTION license, so it is free to use.

What category does Box belong to?

Box is listed under mcp-servers in the Claudeers registry of Claude-compatible tools.

26 views
★ 872 stars
unclaimed
updated about 1 month ago

// embed badge

Box on Claudeers
[![Claudeers](https://claudeers.com/api/badge/box.svg)](https://claudeers.com/box)

// retro hit counter

Box hit counter
[![Hits](https://claudeers.com/api/counter/box.svg)](https://claudeers.com/box)

// reviews

// guestbook

0/500

// related in MCP Servers

🔓

f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete…

// mcp-serversf/⟨HTML⟩★ 172,096◷ NOASSERTION[ claude ]
🔓

A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Gemini CLI & Hermes Agent. Only official website: ccswitch.io

// mcp-serversfarion1231/⟨Rust⟩★ 140,512◷ MIT[ claude ]
🔓

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

// mcp-serversJuliusBrussee/⟨JavaScript⟩★ 107,719◷ MIT[ claude ]
🔓

An open-source AI agent that brings the power of Gemini directly into your terminal.

// mcp-serversgoogle-gemini/⟨TypeScript⟩★ 107,167◷ Apache-2.0[ claude ]

// built by

→ see how Box connects across the ecosystem