← Back to home@LJH-snow

dsh-tool-firecrawl

Firecrawl web scraping tools for DeepSeek Harness

Stars
0
Language
TypeScript
Created
Oct 5, 2026
Updated
Oct 6, 2026

Introduction

dsh-tool-firecrawl

English | 中文

Firecrawl web scraping integration for DeepSeek Harness (dsh) as a Cordis plugin. The agent can scrape pages, map sites, search the web, and manage crawl jobs while the API key and oversized content stay out of tool output.

Install

npm install @libai168/dsh-tool-firecrawl

Requires @deepseek-ai/cordis (^4.0.1) and @deepseek-ai/dsh-tools (^0.1.0-rc.6) as peer dependencies.

Configuration

- name: 'github:LJH-snow/dsh-tool-firecrawl'
  config:
    # baseUrl: 'https://api.firecrawl.dev'
    apiKeyEnv: 'FIRECRAWL_API_KEY'
    # timeoutMs: 120000

The plugin reads the Firecrawl API key from the environment variable named by apiKeyEnv (default: FIRECRAWL_API_KEY). Do not put a usable key in source, examples, tests, or committed configuration. Create the key in the Firecrawl dashboard and grant only the access your deployment needs.

The baseUrl override must be an absolute http:// or https:// root URL. Only publicly reachable hosts are allowed: localhost, loopback, private, link-local, CGNAT, multicast, reserved/documentation/benchmark ranges, and every IANA special-purpose block are rejected, and a hostname whose DNS results contain any such address fails closed before the request is sent. Credentials, query strings, fragments, and non-root paths are not allowed.

Tools

ToolDescriptionWrite
firecrawl_auth_testVerify the API key and report remaining credits without returning the keyNo
firecrawl_scrapeScrape one URL to markdown plus page metadataNo
firecrawl_mapMap the URLs of one site, optionally ranked by a search queryNo
firecrawl_searchSearch the web and return titles, URLs, and descriptionsNo
firecrawl_crawl_startStart a crawl job and return its ID (billable)Yes
firecrawl_crawl_statusRead one crawl job's status and page metadataNo
firecrawl_crawl_cancelCancel one crawl jobYes

Security contract

  • The API key is read from an environment variable at plugin startup and is never included in tool output or rendered text.
  • Scrape markdown is capped at 20,000 characters, descriptions at 500, titles at 300, map links at 200 entries, search results at 100 entries, and crawl status pages at 50 entries.
  • Crawl status returns page metadata only; page content from crawl results is never forwarded.
  • Crawl pagination (next) URLs are accepted only when they stay on the configured Firecrawl API host, so the bearer key is never sent to a foreign host.
  • firecrawl_crawl_start defaults to 20 pages and clamps the page limit at 200 to bound credit usage; it and firecrawl_crawl_cancel are marked kind: 'edit'.
  • The client passes caller cancellation signals through to fetch and uses a 120-second timeout by default.
  • API failures are normalized into { ok: false, reason } or { found: false, reason } tool results.

API scope

This version uses the Firecrawl v2 public API: GET /v2/team/credit-usage, POST /v2/scrape, POST /v2/map, POST /v2/search, and the POST /v2/crawl, GET /v2/crawl/{id}, DELETE /v2/crawl/{id} job lifecycle. Structured extraction, batch scrape, parse, monitors, and browser interactions are intentionally not included yet.

Development

npm install
npm run typecheck
npm test
npm run build
npm pack --dry-run

License

MIT