Open source · Apache-2.0
中文 ↗

Autohand

The open-source browser AI agent: ask about any page, extract data — and, uniquely, transcribe video.

An AI workbench in your sidebar: reads the page, answers questions, structures data into tables — and transcribes video to a transcript in one line. Local-first, your keys, anonymous opt-out telemetry.

★ GitHub

Install / Update

Latest:loading…   Installed:unknown


          
First-install guide
  1. Click ① to create or pick an empty folder (e.g. autohand-ext).
  2. Click ② — the build is written into that folder.
  3. Click ③ and follow the dialog to load or refresh on chrome://extensions.
  4. For future updates, repeat ②③; Chrome usually auto-reloads after files are overwritten.

Chrome / Edge only (requires the File System Access API).

③ Reload the extension

Web pages cannot navigate to chrome:// URLs — copy the address below and paste it into the address bar:

chrome://extensions/

First install (not loaded yet)

  1. Turn on Developer mode (top right of the extensions page).
  2. Click "Load unpacked" and pick the folder from step ①.

Update (already loaded)

  1. Find the Autohand card on the extensions page.
  2. Click ↻ (reload); Chrome usually auto-reloads after step ② anyway.

Setup after install

Two minutes in the sidebar settings — bring your own keys, or go fully local.

1 · Configure the chat model

Open the sidebar → Settings → Models:

  1. Pick a provider preset — OpenAI-compatible or Anthropic-compatible.
  2. Paste your API key (stored locally, sent only to the endpoint you chose).
  3. Self-hosted / local endpoints (Ollama, vLLM…) work too — any OpenAI-compatible URL.

2 · Speech recognition (ASR)

local option

Sidebar → Settings → Models → Speech recognition (ASR):

  • Cloud presets — Groq (free tier), SiliconFlow, or OpenAI; paste the key, done.
  • Local model — switch the provider to Local model, pick a Whisper size and hit Download: tiny ≈40MB (fastest) · base ≈80MB (recommended) · small ≈250MB (best quality).
  • Local runs entirely in your browser (transformers.js) — audio never leaves the device, no key needed.

Then just say “transcribe this video” on any supported page. Full docs: GitHub README

What it can do

Open any page and talk in the sidebar — read, answer, transform, store. All local-first.

💬 Page Q&A & analysis

Answers grounded in the live page: summarize long articles, explain content, compare options. Reads first, never makes things up.

📊 Data extraction & processing

Structure page content into tables (CSV download) and markdown notes — searchable and reusable.

🎬 Video transcription

unique among open source

MSE sites like Bilibili and Douyin, plus direct links and HLS — say "transcribe this video": the audio track is detected, extracted locally, transcribed with resumable ASR into a timestamped markdown transcript.

📄 Full-page export & archive

Export the whole page as markdown (images localized); data, transcripts, and notes live in one searchable Data tab.

$0
local transcription · runs in your browser
33
gated capabilities · per-action authorization
0
content servers / accounts · anonymous opt-out telemetry
40MB+
local ASR model · download once, transcribe offline

Compared with open-source peers

Autohand Nanobrowser browser-use
Page Q&A / understanding ✓ read-before-answer, verified context ✓✓
Video/audio transcription ✓ local extraction + resumable ASR ✗✗
Full-page markdown export (localized images) ✓ ✗✗
Tables → Data tab (CSV) + markdown notes ✓ ✗✗
Per-action network authorization (tool+target bound) ✓ global or none—
Local-first (media never leaves device) ✓ partialmostly cloud browsers
Task reliability (plans / resume / breaker) ✓ basicbasic
Form factor Chrome side-panel extension Chrome extensionPython/TS library

Comparison based on public features as of 2026-09; file an issue if anything is off.

✓ Anonymous telemetry, one-tap off ✓ No servers ✓ No accounts

Media is processed in the browser; agent-initiated network requests need per-action approval. Privacy