Back to home@Godners-Code

dao-zang-skill

No description

Stars
0
Language
Python
Created
Aug 24, 2026
Updated
Aug 24, 2026
GitHub repo

Introduction

dao-zang-skill

DaoZang offline retrieval & original-text extraction skill for DeepSeek Harness (DSH).

Search 285,117 scripture chunks of the Daoist Canon (《中华道藏》《正统道藏》) by keyword or semantics, and extract exact original text from the source Markdown with line numbers and hit markers. Fully offline — no embedding API, no network needed for retrieval.

Install

dsh plugin --profile web add dao-zang-skill

Or from source: dsh plugin --profile web add https://github.com/godners/dao-zang-skill

What you get

  • text engine (default, zero deps): ChromaDB full-text filter + TF/IDF ranking
  • semantic engine (optional): local bge-m3 ONNX model, same 1024-dim cosine vectors as the database
  • original-text extraction (--original): locates the hit in the raw .md with ⟦...⟧ markers and line numbers
  • file filter (--source): restrict search to files whose name contains a keyword
  • one-click workspace setup: downloads data from the Godners/DaoZang dataset (3,152 markdown files + 6 parquet shards with bge-m3 embeddings) and rebuilds the local ChromaDB offline

Data

The workspace needs ChromaDB/ (285,117 chunks) and Markdowns/ (3,152 files). Prepare it with:

python assets/dao-zang/scripts/setup_workspace.py --dir <workspace>

See USAGE.md for details.

License

MIT