
jarvis
J.A.R.V.I.S. voice assistant with an Iron Man holographic UI, driven by Claude Code. Built on adewaskar/jarvis, with Windows support and a loopback-only bridge.
Install with your AI
Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.
Install and set up jarvis (git-clone project) into my current project. Found on https://claudeers.com/jarvis-2 Repo: https://github.com/ehteshambuildagents/jarvis Homepage/docs: — Detected install method: git-clone → git clone https://github.com/ehteshambuildagents/jarvis Category: automation. Platforms: cli, api, desktop, web, mobile. Read the repo's README for exact setup and env vars, then install it and wire it into my project. Claudeers Health Verdict: active; community-verified: false. Confirm the source before running anything.
git clone https://github.com/ehteshambuildagents/jarvis
// compatibility
| Platforms | cli, api, desktop, web, mobile |
|---|---|
| Operating systems | — |
| AI compatibility | claude |
| License | MIT |
| Pricing | open-source |
| Language | TypeScript |
J.A.R.V.I.S.
A browser voice assistant with an Iron Man holographic interface. Say "Hey Jarvis", he wakes, listens, and does real things through your tools — searches the web, generates images, drives your phone, reads your mail. The face is a web page (React + Vite + Three.js + custom GLSL). The brain is Claude Code, run headless as a library.
The only subscription you need is Claude Code. No API keys, no OpenAI account, no cloud bill — the brain runs on your existing Claude Code login, and the heavy work (the model itself) runs on Anthropic's servers, so even a low-end laptop only has to draw the interface. ElevenLabs is an optional add-on that gives JARVIS a much better voice and sharper hearing; without it he speaks and listens through the browser's own speech, and everything still works.
What this fork adds
- Bridge bound to loopback. The bridge used to listen on every network
interface, and its
Origincheck is a browser convention, not authentication — any client on the same network could set the header and drive an agent that may have shell access. It now binds to127.0.0.1only (bridge/server.mjs). - Windows support. The
chrome_*extension bridge does not run on Windows, so onwin32the system prompt points JARVIS at the Chrome DevTools MCP server instead and teaches it the Git Bash quirks of Windows commands (taskkill //IM,winget). macOS keeps the original behaviour. The prompt also requires a spoken confirmation before shutdown, restart, sign-out or deleting files. - "Jarvis, logout". A voice dismissal that acknowledges and returns to
standby. The pattern must match the whole utterance, so "log out of my
account" is still treated as a task rather than a dismissal (
src/App.tsx). - Unattended mode.
?auto=1powers up on load, with no INITIALISE click or clap, so JARVIS can run as a background assistant. - Windows launchers. Start JARVIS in a Chrome window parked off-screen (kept un-throttled so the wake word still works), bring it back into view, and stop exactly the processes it started. See Running on Windows.
Requirements
In one line: a Claude Code subscription, plus two free things every computer can have — Node.js and Chrome. That's the whole list.
- Claude Code, installed and logged in — this is the only account you need.
Install it with the official method —
npm install -g @anthropic-ai/claude-code, or the platform installer at https://docs.claude.com/en/docs/claude-code — then runclaudeonce and complete login. The bridge reuses that login. No API key, and usage is billed to your existing Claude account. - Node.js 20 or newer — free, one installer from https://nodejs.org. This is a Node web app, so it is the one unavoidable tool.
- Google Chrome or Microsoft Edge, in a real browser window — not an embedded preview pane. Preview panes (including the one inside editors and Claude Code) block microphone access, so the page loads and looks right but never hears you. JARVIS also needs WebGL, which these browsers provide.
- Optional: an ElevenLabs API key — a good add-on, not a requirement. It gives a better voice and sharper transcription; the free tier is plenty for a demo. Without it, everything runs on the browser's own speech.
Run npm run setup after cloning and it checks all of this for you, in plain
language.
Quick start
First, install, then start it:
npm install
npm start # runs the brain and the face together
Then open the URL it prints (http://localhost:5173) in Chrome, click INITIALISE, and say “Hey Jarvis”.
Prefer two terminals? Run them separately instead:
npm install
Terminal 1 — the brain:
npm run bridge
Terminal 2 — the face:
npm run dev
Then open the app in a real Chrome or Edge window:
open http://localhost:5173
Click INITIALISE, allow the microphone when asked, and say "Hey Jarvis".
It has to be a real browser window. Embedded preview panes block the microphone, so JARVIS will look perfectly alive and simply never respond.
Running on Windows
Double-click launchers live in the repo root:
| File | Does |
|---|---|
start-jarvis.cmd | Visible mode: starts the bridge and the interface, then opens Chrome |
start-jarvis.vbs | Hidden mode: runs jarvis.ps1, which starts everything and parks the interface in a Chrome window off the edge of the desktop — you just talk |
show-jarvis.vbs | Brings the hidden window back on screen |
stop-jarvis.vbs | Stops only the processes JARVIS started |
Both start modes run with writes enabled — read Enabling actions first.
The hidden launcher uses a Chrome profile of its own (%LOCALAPPDATA%\JarvisApp)
with the microphone pre-allowed for localhost:5173 and the camera denied,
because nobody can answer a permission prompt in a window they cannot see.
For the ElevenLabs voice, copy jarvis-secrets.example.cmd to
jarvis-secrets.cmd and paste your key in. That file is git-ignored.
For browser control, add the Chrome DevTools MCP server to your user-level Claude Code config, where the bridge looks for it:
claude mcp add --scope user chrome-devtools -- npx chrome-devtools-mcp@latest
How it works
JARVIS is two processes. The browser is the face and the voice; the bridge is the brain and the hands.
┌─ browser (the face) ───────────────┐ ┌─ bridge (the brain) ─────────────┐
│ "Hey Jarvis" wake word │ │ Node · bridge/server.mjs │
│ local VAD → speech to text │ ws │ Claude Agent SDK │
│ reactor UI (Three.js + GLSL) │◄─────► │ = Claude Code, headless │
│ text to speech │ 8787 │ spawns your MCP servers │
│ heads-up display │ │ permission gate (decideTool) │
└────────────────────────────────────┘ └──────────────────────────────────┘
Everything you see and hear happens in the browser. The bridge is a single Node
process (bridge/server.mjs) that runs the Claude Agent SDK
(@anthropic-ai/claude-agent-sdk) — this spawns the real claude CLI as a child
process, so the brain literally is Claude Code, headless. They talk over a
WebSocket (plus a few HTTP endpoints) on ws://localhost:8787.
Why a bridge at all? A browser tab cannot spawn the local stdio MCP servers —
higgsfield, elevenlabs, android, playwright, exa, serper, and the
rest. The bridge can. And because it is the Agent SDK, it authenticates off your
existing Claude Code login: no API key, billed to that same Claude account.
The model. claude-opus-5 at effort medium by default. Override with the
JARVIS_MODEL and JARVIS_EFFORT environment variables. On startup the bridge
prints its choice, e.g. [jarvis] model claude-opus-5 · effort medium.
The voice pipeline
The loop is designed so that nothing silently dies and barge-in feels natural.
- Detection is local. An energy-based voice-activity detector
(
src/lib/vad.ts) decides when you are speaking. It is instant, cannot quietly fail, and is what makes barge-in work — speak while JARVIS is talking and he stops. - Transcription has two tiers, chosen automatically at boot. The browser asks
the bridge
/healthand picks the best available:- ElevenLabs key present → ElevenLabs Scribe, via the bridge
/sttendpoint. - Nothing configured → the browser's own
SpeechRecognition(Chrome/Edge), guarded by a heartbeat so it recovers when Chrome throttles it.
- ElevenLabs key present → ElevenLabs Scribe, via the bridge
- Speaking uses the ElevenLabs voice when a key is present, and the
browser's
speechSynthesisotherwise. If a cloud call fails it falls back to the browser voice, and if the OS voice itself is broken it latches over to the cloud voice.
So it works with no keys and auto-upgrades when a key appears — there is no flag
to set. Capability detection lives in src/lib/capabilities.ts, which probes the
bridge's GET /health (returning { ok, tts, stt }, both tracking the
ElevenLabs key) once at boot and picks the engines.
What JARVIS can do
Beyond answering, JARVIS reaches every MCP server in your Claude Code configuration, and can drive his own interface.
Your tools
Every server in your ~/.claude.json is handed to the SDK explicitly. Depending
on what you have installed, that is roughly:
- Web & search —
exa,serper,serpapi - Images & video —
higgsfield,openrouter-image,palmier-pro - Voice —
elevenlabs - Your phone —
android - The browser —
playwright
A few things you can say:
- "What's happening in AI this week?"
- "Generate an image of the Mark VII suit."
- "Take a screenshot of my phone."
- "Open my GitHub notifications."
Note on account connectors. Servers you added through your claude.ai account are not stored on disk, so the bridge cannot see them — it works from the servers in
~/.claude.json(about 14), not the claude.ai ones.
JARVIS controls the interface
He drives the UI through MCP tools the bridge exposes:
ui_theme— accent, background, per-phase coloursui_reactor— colour, scale, intensity, spin, and style (ring|sphere|wire), visibilityui_orbit— put images in orbit around the reactorui_chrome— show or hide rails, transcript, badgesui_effect—glitch|pulse|scan|shake|flashui_screen— clearui_reset— back to defaults
So "make it red, hide the systems list, put that render in orbit" is a spoken command.
The heads-up display
JARVIS authors panels with a display tool against a fixed .hud-* design
system. The browser sanitises the markup (DOMPurify, a class allowlist and a
strict CSP) before rendering. Rich media works — images, <video>, and
YouTube/Vimeo embeds. Remote images and video are fetched server-side through
the bridge (/img and /media, both SSRF-guarded), so hotlink-blocked news
thumbnails still appear and the page never beacons your IP to a host the model
chose.
Controls
| Key / phrase | Does |
|---|---|
| "Hey Jarvis" | Wake him |
| Space | Talk without the wake word |
| Just speak | Interrupt him mid-sentence (barge-in) |
| V | Cycle the browser voice |
| "Jarvis, logout" | Acknowledge and stand down (also "stand down", "goodbye", "that'll be all") |
| Escape | Stand down |
| D | Live diagnostics panel |
| T | One-line audio self-test |
The boot sequence
Power-up plays a four-beat Iron Man start-up (src/ui/Boot.tsx): an
"INITIATING SYSTEM" status bar with a segmented progress bar and boot log; then
concentric reticle rings resolving into "J.A.R.V.I.S"; then a suit schematic;
then the triangular arc reactor lighting up — with a start-up sound under it
(public/audio/boot-music.mp3).
Configuration
Everything is optional in bridge mode. Frontend settings live in .env.local
(copy .env.example); bridge settings are environment variables.
Bridge
| Variable | Default | Effect |
|---|---|---|
JARVIS_BRIDGE_PORT | 8787 | Port for the WebSocket + HTTP endpoints |
JARVIS_MODEL | claude-opus-5 | Model to run |
JARVIS_EFFORT | medium | Reasoning effort |
JARVIS_ALLOW_WRITES | off | 1 allows effectful tools (see below) |
JARVIS_ALLOWED_ORIGINS | local dev | Extra WebSocket origins to accept |
JARVIS_ALLOW_NO_ORIGIN | off | Accept connections with no Origin header |
JARVIS_FILE_ROOTS | — | Roots the /file endpoint may serve from |
JARVIS_VOICE_ID | — | ElevenLabs voice id |
ELEVENLABS_API_KEY | — | Optional; enables the ElevenLabs voice + Scribe |
Frontend (.env.local)
| Variable | Effect |
|---|---|
VITE_BACKEND | bridge (default) or direct |
VITE_BRIDGE_URL | Where to reach the bridge |
VITE_TTS_ENGINE | system or kokoro |
VITE_KOKORO_VOICE | Voice for the Kokoro engine |
VITE_USE_ELEVENLABS | Force the ElevenLabs voice on |
VITE_ANTHROPIC_API_KEY | Direct mode only |
Adding an ElevenLabs key
You do not have to touch a flag. Either:
- Set
ELEVENLABS_API_KEYon the bridge before starting it, or - Add the key to your
elevenlabsMCP server's env in~/.claude.json— the bridge reads it from there too.
Either way, /health starts reporting the capability, the browser picks it up on
the next boot, and both the voice and transcription upgrade automatically.
Enabling actions
The tool gate starts read-only. Search, generation and lookups run freely;
anything effectful — send, tap, delete, install, pay — is denied. Voice is a poor
interface for a confirmation dialog, so the decision is made ahead of time in
decideTool() in bridge/server.mjs, not at the moment of use. The bridge sets
settingSources: [], which makes its own gate the only authority — filesystem
settings and any global bypassPermissions cannot override it.
To allow effectful tools (phone, browser driving, sending), run the bridge this way instead:
npm run bridge:writes
Read
decideTool()before you do. "Hey Jarvis, clean up my downloads folder" means something rather different with writes enabled.
Troubleshooting
I can't hear him, or he can't hear me. Press D for the diagnostics panel — it states plainly whether he is hearing you and whether he is producing sound. Press T for a one-line audio self-test.
No voice at all. You must be in Chrome or Edge, in a real browser window (not an embedded preview), and you must have allowed the microphone.
Bridge not reachable. Check that npm run bridge is still running in its
terminal, and that nothing else is holding port 8787.
Security
All of this lives in bridge/server.mjs:
- The bridge listens on
127.0.0.1only, so nothing else on the network can reach it. The origin check below is a browser convention, not authentication. - The WebSocket accepts only local dev origins (add more with
JARVIS_ALLOWED_ORIGINS). /file,/imgand/mediavalidate the scheme, confine to allowed roots, resolve the real path, and refuse private and loopback addresses (SSRF guard).- The tool gate (
decideTool) is default-deny for effectful MCP tools. - A strict CSP in
index.html; model-authored panel HTML is sanitised.
Credits & licence
MIT. The original project is by Aditya Dewaskar — adewaskar/jarvis. The additions in this fork are released under the same licence.
Music by Kevin MacLeod (incompetech.com), licensed under Creative Commons
Attribution 4.0 — see public/audio/CREDITS.md.
The boot sound and any tracks in public/audio/ ship with the project for the
demo. If you go on to monetise something built on this, clearing the rights to
that audio is your responsibility.
// faq
What is jarvis?
J.A.R.V.I.S. voice assistant with an Iron Man holographic UI, driven by Claude Code. Built on adewaskar/jarvis, with Windows support and a loopback-only bridge.. It is open-source on GitHub.
Is jarvis free to use?
jarvis is open-source under the MIT license, so it is free to use.
What category does jarvis belong to?
jarvis is listed under automation in the Claudeers registry of Claude-compatible tools.
// embed badge
[](https://claudeers.com/jarvis-2)
// retro hit counter
[](https://claudeers.com/jarvis-2)
// reviews
// guestbook
// related in Automation & Workflows
The API to search, scrape, and interact with the web at scale. 🔥
🌐 Make websites accessible for AI agents. Automate tasks online with ease.
Taste-Skill - gives your AI good taste. stops the AI from generating boring, generic slop