Trawl

CLI

The trawl CLI is the fastest way to manage scraps from your terminal. Install once, authenticate once, then list, run and debug from anywhere — no browser needed. Pairs well with CI/CD pipelines and the Claude Code skills.

Install

cURL
npm install -g @trawlme/cli

Or run ad-hoc without installing:

cURL
npx @trawlme/cli list

Authenticate

Interactive (default)

cURL
trawl login

Prompts for email and password. Credentials are stored locally for subsequent commands.

Interactive trawl login also installs any bundled Claude Code skill you don't already have, and tells you to restart Claude Code to see it — skills load at session start, so anything installed mid-session stays invisible until then. Opt out with TRAWL_SKILLS_SYNC=0. This only happens on interactive login — every other command only ever suggests running trawl skills install, never installs anything itself. See Skills.

CI/CD — environment variable

cURL
export TRAWL_TOKEN=<your-jwt>
trawl list

TRAWL_TOKEN overrides the stored session for the current command — no trawl login call needed, and nothing is written to disk. Ideal for GitHub Actions, Vercel, Netlify, and any other non-interactive context.

Requires @trawlme/cli ≥ 1.17.0. On earlier versions, TRAWL_TOKEN is only consumed by trawl login itself (to persist the token) — run trawl login --token <jwt> first, then subsequent commands read the stored session.

CI/CD — token flag

cURL
trawl login --token <jwt>

Custom API URL (self-hosted)

cURL
trawl login --url https://self-hosted.example.com

Common workflows

The verbs you reach for daily — list, get, run, trigger, data, history, run-info — sit at the top level. Each also still answers under trawl scraps, kept working for older scripts, but only the short form shows up in trawl --help, so that is the one used throughout this guide.

Everything else you do to a scrap — editing it, deleting it, its credentials, snapshots, banner, doctor, autofix, watching a run — stays under trawl scraps and has no short form. One name belongs to both: trawl create is the AI wizard that writes a scrap for you, while trawl scraps create takes a script you already have. Different commands, not two spellings of one.

List your scraps

cURL
trawl list
trawl list --json                        # machine-readable JSON
trawl list --status failure              # filter by last run status
trawl list --status success
trawl list --status never                # scraps that have never run
trawl list --unhealthy                   # only scraps whose schedule is paused for failing

list and get show a health column next to the status. It reads ⚠ cron paused when the scrap's last scheduled run was skipped because the scrap keeps failing, and — otherwise. That column is a breadcrumb of the last tick, so it can lag by one run and is never set on a scrap with no schedule.

--unhealthy asks the server for the live answer instead: only scraps that failed their last 3 runs in a row, each annotated with how long the streak is and when it started. Under --json those two fields are present only when you pass --unhealthy. A server older than the feature ignores the filter and returns everything — the CLI notices, warns on stderr, and wraps the --json payload as { "scraps": [...], "_healthFilterUnconfirmed": true } so a script can tell the difference. Requires @trawlme/cli ≥ 3.9.0.

See Schedule for what the paused state means and how a scrap recovers from it.

Get a single scrap

cURL
trawl get <id>
trawl get <id> --json

Run a scrap on demand (UI / JWT)

cURL
trawl run <id>           # blocks until the run completes
trawl run <id> --watch   # blocks, streaming activity output as it runs

run calls GET /api/scraps/load/:id authenticated with your JWT (interactive login) and blocks for the whole synchronous run (typically 30–250s) — it does not return until the scrap finishes or errors. It is the same code path used by the Trawl web interface and is rate-limited per user. For a non-blocking, fire-and-forget launch, use trigger below.

Trigger a scrap without blocking (CI/CD)

cURL
trawl trigger <id>           # async: queues the run, returns immediately
trawl trigger <id> --wait    # synchronous: block until the run completes
trawl trigger <id> --watch   # queue, then stream the activity log until the run ends

trigger calls POST /api/scraps/worker/:id and accepts either credential the CLI supports: a scoped TRAWL_API_KEY (recommended for CI/CD — revocable on its own, independent of any human's login, and not subject to the 7-day expiry a session JWT carries) or a session JWT via TRAWL_TOKEN/a stored trawl login. TRAWL_API_KEY wins when both happen to be set. The difference from run is the endpoint's behavior, not the auth method.

By default trigger is asynchronous — it queues the worker run and returns immediately instead of holding the connection for the whole 30-250s run. For pipelines that must gate on the result, add --wait (block until completion) or --watch (stream until the run ends); both make the command's exit reflect the run's outcome.

CI/CD tip: use trigger (not run) in GitHub Actions, Vercel hooks, and other automated contexts — it won't hold your job's connection open for the whole run. Add --wait or --watch when the step must block on the result.

yaml
# GitHub Actions example (requires @trawlme/cli >= 3.8.0 for TRAWL_API_KEY;
# TRAWL_TOKEN works from >= 1.17.0)
- name: Trigger scrap
  env:
    TRAWL_API_KEY: ${{ secrets.TRAWL_API_KEY }}
  run: npx @trawlme/cli trigger ${{ vars.SCRAP_ID }} --watch

Watch a run live (SSE stream)

cURL
trawl scraps watch <id>

Attaches to the SSE stream for the scrap's next run. Use after run or trigger to follow output in real time.

Get recent results

cURL
trawl data <id>          # latest run payload
trawl data <id> --json   # raw JSON
trawl data <id> --fresh  # run it now and return that — spends a run from your quota

Debug a failing scrap

When a scrap returns errors, empty results, or has been auto-fixed, use the diagnostic commands to understand what happened without opening the web UI.

Quick diagnosis — last run status + failed selector + autofix summary:

cURL
trawl scraps doctor <id>
trawl scraps doctor <id> --json   # raw run + autofix JSON

doctor fetches the last run and prints: status badge, error message, failed CSS selector, block detection status, proxy tier (Tier 0–4 abstract label), empty-context page URL/anchor count, and a one-line autofix summary if the engine ran.

Cost, proxy vendor, and egress IP are not exposed — by design.

When the run landed on a login page, doctor shows a dedicated hint instead of the generic error line:

  Login wall (hypothesis): this looks like a login redirect — your own session cookies are the likely fix.
  → trawl scraps account session capture <id>
    headless/CI: trawl scraps account session set <id> -c cookies.json
    or app Settings → Account

Requires @trawlme/cli ≥ 3.12.3 (npm view @trawlme/cli dist-tags shows which tag carries it). Versions before 3.9.1 show no login-wall hint at all; 3.9.1 to 3.12.2 show it led by session set.

This is a hypothesis, not a promise — capturing a session is the right lever to pull, not a guaranteed fix on every site. See Account & sessions.

List past runs + inspect a specific one (any run, not just the last):

cURL
trawl history <id>            # table of recent runs: status, time, proxy tier, failureKind
trawl history <id> --json     # newest first; --limit <n> caps rows (default 20)
trawl run-info <hid>          # one run's detail by its history id (status, tier, failureKind, error + selector)
trawl run-info <hid> --json

doctor/autofix inspect the last run; history + run-info reach any run. List runs with history, take a run's hid, then pass it to run-info for the full diagnosis. run-info returns the proxy tier + failureKind to the scrap owner via GET /api/historys/:hid. failureKind includes 'auth' for a run that landed on a login page — run-info --json (along with doctor --json and data --errors --json) carries a docs field pointing at the account-sessions guide whenever the run's live verdict is 'auth'.

Full autofix detail — decision, diff, dry-run results, knowledge used:

cURL
trawl scraps autofix <id>
trawl scraps autofix <id> --json

autofix shows the fix decision (applied / failed / skipped), classification, the unified diff of the code change, dry-run pass rate, and the knowledge fingerprints the engine consulted.

Failure diagnostics inline with data:

cURL
trawl data <id> --errors   # diagnosis when the last run failed; item count otherwise

Download the captured page HTML (to inspect what the worker actually saw):

cURL
trawl scraps snapshot <id>                     # captured page at end of run
trawl scraps snapshot <id> --error             # error-path snapshot (selector failure context)
trawl scraps snapshot <id> --error -o page.html  # write to file instead of stdout

Typical debug flow:

cURL
trawl scraps doctor <id>           # 1. what went wrong + autofix outcome
trawl scraps autofix <id>          # 2. if autofix ran — inspect the diff + dry-run
trawl scraps snapshot <id> --error # 3. inspect the DOM the worker saw when it failed
trawl run <id> --watch             # 4. trigger a re-run after fixing the script

Create or update a scrap

cURL
trawl scraps create -t "HN front page" -u https://news.ycombinator.com
trawl scraps update <id> --cron "0 * * * *" --alert me@example.com
trawl scraps update <id> --no-cron   # disable schedule

Manage account credentials (for auth-gated scraps)

cURL
trawl scraps account set <id> -u username -p password
trawl scraps account status <id>
trawl scraps account clear-session <id>  # force a fresh login next run
trawl scraps account delete <id>
trawl scraps account session set <id> -c <file>       # upload a session from a file
trawl scraps account session capture <id> [--chrome <path>]  # capture a session interactively

account session set -c <file> accepts either a bare Puppeteer cookie JSON array, or a { cookies, origins } storageState-shaped file — the same shape session capture produces.

account session capture <id> opens a real, visible Chrome window at the scrap's target URL, waits for you to log in there — 2FA included — then reads the resulting session over the Chrome DevTools Protocol (cookies and per-origin local storage) and uploads it.

You stay authenticated as yourself throughout — Trawl never sees your credentials, only the resulting session. Responsibility for lawful use of that session stays with you; this isn't legal advice.

Requires a local interactive terminal with a display — it does not work headless, in CI, or over a plain SSH session. Auto-detects a system Chrome/Chromium; override with --chrome <path>, TRAWL_CHROME_PATH, or the common PUPPETEER_EXECUTABLE_PATH/CHROME_PATH conventions. Completion is explicit: press Enter in the terminal once you're signed in. Closing the Chrome window also completes the capture, but cookies only (not local storage), and only when Chrome itself survives that close — on a setup where closing the last window quits Chrome entirely, nothing is captured and the command says so; press Enter there instead.

The capture is scoped to the scrap's own target host: a cookie explicitly scoped to a domain is in scope for that host and its subdomains, while a host-only cookie (no Domain attribute) is in scope only for the exact host it was set on — a host-only cookie set on a sibling host is excluded, since a real browser would never send it to the target host either.

Both session set and session capture refuse a plaintext (non-https) API URL — a session is a bearer-equivalent secret. Loopback hosts (127.0.0.1, ::1, localhost) are exempt, for local/self-hosted testing.

--json on session capture never prints the session itself — only counts:

JSON
{
  "account": { "hasSession": true, "sessionSavedAt": "..." },
  "targetDomain": "example.com",
  "capture": {
    "cookiesCaptured": 4,
    "cookiesDroppedOutOfScope": 1,
    "cookiesDroppedInvalid": 0,
    "originsCaptured": 1,
    "originsDroppedOutOfScope": 0,
    "originsUnreadable": 0,
    "closedEarly": false
  }
}

See Account & sessions for the full auth-gated scraping workflow and session model.

Delete a scrap

cURL
trawl scraps rm <id>
trawl scraps rm <id> --force   # skip confirmation prompt

Install Claude Code skills

The CLI bundles the @trawlme/skills package. Install all Trawl skills into Claude Code in one step:

cURL
trawl skills install          # → ~/.claude/skills/ (user-level, all projects)
trawl skills install --local  # → .claude/skills/ (project-level)
trawl skills list             # show bundled skills and install status
trawl skills update           # reinstall to pick up CLI version bumps

trawl login also installs any skill you're missing automatically — see Authenticate above. See Skills for what each skill does.

Discovering the docs

cURL
trawl spec --json

Prints a versioned, machine-readable description of the whole command tree — useful for a script or an agent driving this CLI without hand-maintained knowledge of it. Among other fields, it carries docsUrl and llmsUrl (the guide root and its plain-text digest, when the CLI can resolve them for the configured API), and each scraps account * command carries its own docs deep link, straight to this guide's Account & sessions companion. A run whose failureKind is 'auth' carries that same docs link in doctor/data --errors/run-info's JSON payload (see Debug a failing scrap above).

CI/CD usage

Use trigger (not run) in automated contexts — it's async by default and won't hold your job's connection open for the whole run. Both run and trigger accept either credential the CLI supports: a scoped TRAWL_API_KEY (recommended — revocable independently of any human's session, and not tied to the 7-day session-JWT cycle: keys default to a 1-year expiry, up to 2 years) or a session JWT, supplied via TRAWL_TOKEN (env var — requires @trawlme/cli ≥ 1.17.0) or trawl login --token <jwt>. TRAWL_API_KEY takes precedence when both are set.

A session JWT expires after 7 days — if you use the JWT fallback instead of an API key, plan to refresh the stored secret periodically, e.g. by running trawl login && trawl token and updating it.

GitHub Actions

yaml
- name: Trigger scrap
  env:
    TRAWL_API_KEY: ${{ secrets.TRAWL_API_KEY }}
  run: |
    npx @trawlme/cli trigger ${{ vars.SCRAP_ID }} --watch

Create the key from Developers → API Keys and store it as a repository secret named TRAWL_API_KEY. The --watch flag blocks until the run completes, letting you use the exit code for gating.

Vercel / Netlify

Set TRAWL_TOKEN (a session JWT, obtained via trawl login + trawl token) — or TRAWL_API_KEY, recommended, see above — in environment variables, then call npx @trawlme/cli trigger <id> in your build or deploy hooks.

Environment variables

Variable Description
TRAWL_TOKEN Session JWT — overrides the stored session for the current command, no trawl login call needed (requires @trawlme/cli ≥ 1.17.0). 7-day expiry, same as any trawl login session.
TRAWL_API_URL Override the default API URL (https://api.trawl.me)

Command reference

Command Description
trawl login [--url] [--token] [--email] [--password] Authenticate
trawl logout Remove stored credentials
trawl token Print the stored JWT — this is what feeds TRAWL_TOKEN in CI
trawl list [--json] [--status] List scraps
trawl get <id> [--json] Fetch one scrap
trawl create <url> --prompt <goal> [--no-autofix] [--json] Generate a scrap from a URL and a plain-language goal (AI)
trawl scraps create -t <title> [-u <url>] [-r <request>] [-d <description>] [--tier <tier>] Create scrap
trawl scraps update <id> [options] Update scrap
trawl run <id> [--watch] Run via JWT (UI path, GET /api/scraps/load/:id) — blocks until the run completes
trawl trigger <id> [--wait] [--watch] Trigger via JWT (TRAWL_TOKEN or trawl login), CI/CD path, POST /api/scraps/worker/:id; async by default, --wait blocks until done. (trawl scraps trigger <id> still resolves — hidden legacy alias.)
trawl scraps watch <id> Stream live output
trawl data <id> [--json] [--errors] [--fresh] Get latest results; --errors shows failure diagnostics instead when the last run failed; --fresh launches a new run rather than reading the stored one, and costs a run
trawl history <id> [--json] [--limit <n>] List a scrap's past runs (status, time, tier, failureKind), newest first
trawl run-info <hid> [--json] One run's detail by its history id (status, tier, failureKind, error + selector)
trawl scraps doctor <id> [--json] [--autofix] Diagnose last run: error, failed selector, block status, proxy tier, autofix outcome
trawl scraps autofix <id> [--json] Full autofix detail: decision, diff, dry-run results, knowledge used
trawl scraps snapshot <id> [--error] [-o <file>] Download captured page HTML; --error fetches the error-path snapshot
trawl scraps rm <id> [--force] Delete scrap
trawl scraps account set <id> Set account credentials
trawl scraps account status <id> Show account + session state
trawl scraps account clear-session <id> Clear cached session
trawl scraps account delete <id> Remove account credentials
trawl scraps account session set <id> -c <file> Upload a session (cookie array, or { cookies, origins } storageState file)
trawl scraps account session capture <id> [--chrome <path>] Open a headed Chrome, log in, capture the session over CDP, and upload it
trawl spec [--json] Print the machine-readable command tree (docsUrl, llmsUrl, per-command docs links)
trawl skills list List bundled Claude Code skills
trawl skills install [skill] [--local] Install one or all skills
trawl skills uninstall [skill] [--local] Remove skills
trawl skills update [skill] [--local] Reinstall (sync with CLI version)

See also: API keys · Skills · Using Trawl with Claude

Next step → Skills