Back to home

ralfsqual

dsh-pdf-reader

dsh-pdf-reader: DeepSeek Harness plugin adding a read_pdf tool (local PDF text extraction via pdfjs-dist)

Stars
0
Language
JavaScript
Created
Aug 15, 2026
Updated
Aug 15, 2026

Introduction

dsh-pdf-reader

A DeepSeek Harness plugin that adds a read_pdf tool, extracting text from PDF files page by page. Fully local parsing — no upload, no network access.

为 DeepSeek Harness 增加 read_pdf 工具,逐页提取 PDF 文本。纯本地解析,不上传、不联网。

Install / 安装

dsh plugin --profile web add github:ralfsqual/dsh-pdf-reader

Or install from a local checkout / 或本地安装:

dsh plugin --profile web add /path/to/dsh-pdf-reader

Restart dsh web afterwards. The read_pdf tool becomes available to the agent in new sessions.

重启 dsh web 后生效,新会话中 Agent 即可调用 read_pdf

Usage / 使用

Tell the agent to read a PDF, or combine it with an @ file mention (e.g. via dsh-at-file):

Read @docs/spec.pdf and summarize the requirements.
读取 @docs/report.pdf 并总结要点。

Tool parameters

ParameterTypeDescription
pathstringPDF path, relative to the current working directory or an absolute path inside the workspace. / 相对当前工作目录或工作区内绝对路径
maxPagesnumberMax pages to extract, default 500; 0 = no limit. / 最多提取页数
maxCharsnumberMax characters to return, default 120000. / 最多返回字符数

How it works / 工作原理

  • Registered as a model tool via defineTool (@deepseek-ai/dsh-tools).
  • Parses the PDF with pdfjs-dist (2.6.347), extracting text per page with --- page N --- separators.
  • Output is a structured object (pages, extractedPages, truncated, empty, text) with a readable render.

Security / 安全

  • Workspace-confined: path must resolve inside the current working directory; escape attempts (e.g. ../) are rejected. / 路径限定工作区内,越界拒绝。
  • Read-only: the tool only reads the file, never writes, never executes commands. / 只读,绝不写入或执行命令。
  • Local only: parsing happens entirely on the host; no network requests are made. / 纯本地解析,无任何网络请求。
  • Password-protected, corrupted, or non-PDF files produce clear error messages.

Prerequisites / 环境要求

  • DeepSeek Harness with a web (or any agent-capable) profile.
  • Node.js >= 22.19.
  • Works on Windows, macOS, and Linux (host process only; no browser/native code).

Limitations / 已知限制

  • Extracts text layers only. Scanned/image-only PDFs contain no extractable text (the tool reports empty; OCR is out of scope). / 仅提取文本层,扫描件需 OCR。
  • Paths must not contain a leading @ when combined with @-mention plugins. / 与 @ 引用插件合用时路径不能以 @ 开头。
  • Very large PDFs are bounded by maxPages/maxChars; output is truncated with a truncated: true marker.

License

MIT — see LICENSE.