Sep 17, 2026
Turn a Technical Book into an Agent Skill with book-to-skill
Convert a PDF, EPUB, or docs folder into a chapter-indexed skill your agent loads on demand — the README measures 24× to 51× fewer tokens per question than dumping the book into context.
You read a great technical book once, and three months later you cannot remember chapter 7 exists. book-to-skill, a 31k-star converter, turns the book into a structured skill your agent loads on demand — so asking beats digging through the PDF again.
Why This Skill Matters
Run /book-to-skill on a file, a folder, or a glob and it builds a skill in ~/.agents/skills/<slug>/: a SKILL.md with core mental models and a chapter index (about 4,000 tokens), one markdown file per chapter (about 1,000 tokens each), plus a glossary, a patterns file, and a cheatsheet. Chapter files load only when you ask about that topic, so the standing token cost stays small.
Under the hood there are two halves: a deterministic Python extractor turns the document into clean text, and a spec-driven generator — your agent following SKILL.md — restructures that into the skill. The README reports 24× to 51× fewer tokens per question than feeding the whole book to context, measured on real books; that is the project's own benchmark, so treat it as directional rather than a guarantee for your library.
Installation
The one-command route works on any host that reads the Agent Skills standard:
npx skills add virgiliojr94/book-to-skill
Or clone it manually into your skills folder — ~/.claude/skills/book-to-skill for Claude Code, ~/.copilot/skills/ for Copilot CLI, ~/.agents/skills/ for the cross-agent directory.
Real Workflow: Make a Book Queryable
Convert your own copy, passing an optional skill name:
/book-to-skill ./designing-data-intensive-applications.pdf ddia
Before extraction, the skill asks whether the book is technical or text-heavy and picks the tool: docling for code, tables, and formulas (about 1.5 seconds per page), pdftotext for plain prose. A scanned PDF has no text layer to extract — the extractor checks the first pages, stops, and tells you; run ocrmypdf first, then convert.
Once generated, ask about one chapter and the agent reads only that file:
/ddia replication
Real Workflow: Fold Your Team Docs into One Skill
The name says book, but any structured prose works. Point it at your docs/ folder — architecture decision records, runbooks, onboarding guides — and the same extraction produces one skill your team queries while coding. When new material lands, the update/fold-in mode merges it into the existing skill instead of starting over.
Tips
- Check your extractors first with
python3 scripts/extract.py --check— it prints what is installed per format and the exact command to install anything missing. - Formats covered: PDF, EPUB, DOCX, HTML, and RTF, with plain text and Markdown built in; MOBI and AZW go through Calibre's
ebook-convert. - After a conversion, the converter can publish the skill to GitHub (private by default) so your other hosts install it with
npx skills add. - The copyright posture is explicit in the README: processing is local, the output is synthesized notes rather than raw passages, and skills made from copyrighted books should stay private.
When Not to Use This
A quick term lookup does not justify a conversion — plain conversation is faster. You will not get verbatim text back: the skill explicitly never copies raw passages, it produces structured notes, so quotations need the book itself. And do not redistribute generated skills of copyrighted books — the README's own guidance is to keep them private.
See the leaderboard for more skills.