Sep 17, 2026

Turn a Technical Book into an Agent Skill with book-to-skill

Convert a PDF, EPUB, or docs folder into a chapter-indexed skill your agent loads on demand — the README measures 24× to 51× fewer tokens per question than dumping the book into context.

#tutorial#skill-creation#productivity

You read a great technical book once, and three months later you cannot remember chapter 7 exists. book-to-skill, a 31k-star converter, turns the book into a structured skill your agent loads on demand — so asking beats digging through the PDF again.

Why This Skill Matters

Run /book-to-skill on a file, a folder, or a glob and it builds a skill in ~/.agents/skills/<slug>/: a SKILL.md with core mental models and a chapter index (about 4,000 tokens), one markdown file per chapter (about 1,000 tokens each), plus a glossary, a patterns file, and a cheatsheet. Chapter files load only when you ask about that topic, so the standing token cost stays small.

Under the hood there are two halves: a deterministic Python extractor turns the document into clean text, and a spec-driven generator — your agent following SKILL.md — restructures that into the skill. The README reports 24× to 51× fewer tokens per question than feeding the whole book to context, measured on real books; that is the project's own benchmark, so treat it as directional rather than a guarantee for your library.

Installation

The one-command route works on any host that reads the Agent Skills standard:

npx skills add virgiliojr94/book-to-skill

Or clone it manually into your skills folder — ~/.claude/skills/book-to-skill for Claude Code, ~/.copilot/skills/ for Copilot CLI, ~/.agents/skills/ for the cross-agent directory.

Real Workflow: Make a Book Queryable

Convert your own copy, passing an optional skill name:

/book-to-skill ./designing-data-intensive-applications.pdf ddia

Before extraction, the skill asks whether the book is technical or text-heavy and picks the tool: docling for code, tables, and formulas (about 1.5 seconds per page), pdftotext for plain prose. A scanned PDF has no text layer to extract — the extractor checks the first pages, stops, and tells you; run ocrmypdf first, then convert.

Once generated, ask about one chapter and the agent reads only that file:

/ddia replication

Real Workflow: Fold Your Team Docs into One Skill

The name says book, but any structured prose works. Point it at your docs/ folder — architecture decision records, runbooks, onboarding guides — and the same extraction produces one skill your team queries while coding. When new material lands, the update/fold-in mode merges it into the existing skill instead of starting over.

Tips

  • Check your extractors first with python3 scripts/extract.py --check — it prints what is installed per format and the exact command to install anything missing.
  • Formats covered: PDF, EPUB, DOCX, HTML, and RTF, with plain text and Markdown built in; MOBI and AZW go through Calibre's ebook-convert.
  • After a conversion, the converter can publish the skill to GitHub (private by default) so your other hosts install it with npx skills add.
  • The copyright posture is explicit in the README: processing is local, the output is synthesized notes rather than raw passages, and skills made from copyrighted books should stay private.

When Not to Use This

A quick term lookup does not justify a conversion — plain conversation is faster. You will not get verbatim text back: the skill explicitly never copies raw passages, it produces structured notes, so quotations need the book itself. And do not redistribute generated skills of copyrighted books — the README's own guidance is to keep them private.


See the leaderboard for more skills.