Heading contracts, not freeform chat
Every mode writes the same markdown sections so the side panel can parse, stream, and repair output. No unnamed blobs of prose.
A production-grade AI summarizer engineered from first principles in plain JavaScript. Designed as a structured cognitive workspace with PDF and academic-paper extraction, a live multi-phase stepper, visual concept trees and timeline rails, prompt envelopes, quality-gate self-healing, and per-tab isolation.
Choose a reading scenario, then watch the side panel adapt source, mode, and depth. This page does not call model APIs — it mirrors the real workflow with curated sample output.
Each capability exists to remove a common source of friction: copying a transcript by hand, losing timestamps, getting a three-bullet TLDR, or leaking private pages to a cloud chat tab.
Every mode writes the same markdown sections so the side panel can parse, stream, and repair output. No unnamed blobs of prose.
arXiv, PubMed, OpenReview, IEEE, and Chrome PDF text layers extract as first-class sources with an academic prompt persona.
Extract → Analyze → Synthesize → Quality replaces a static spinner. Long papers show chunk counters such as 2/3.
Concepts mode renders filterable Core / Important / Supporting cards. Timeline mode and YouTube details become a vertical milestone spine.
Deep and Long runs score section depth, then regenerate only the failing headings — keeping healthy content and cutting token waste.
Gemini and OpenAI-compatible endpoints sit next to Ollama and LM Studio. Providers receive one prompt string and stay UI-agnostic.
Generic chat leaves you to invent the workflow. Basic extensions stop at a shallow TLDR. DeepDigest keeps extraction, structure, repair, and follow-up in one side panel.
| Need | Generic AI chat in a new tab | Basic summary extensions | DeepDigest |
|---|---|---|---|
| Stay with the source | Copy, paste, context-switch | Popup that covers the page | Chrome side panel beside the tab |
| YouTube transcripts | External grabber or none | Unstructured dump | Native timestamps, chapters, chunking |
| PDFs & academic papers | Copy abstract by hand | Usually unsupported | arXiv, PubMed, OpenReview, IEEE, PDF.js layers |
| Output consistency | Depends on the prompt you wrote | Three-bullet TLDR | Heading contract + quality gate |
| Long-form sources | Hits the context window | Silent truncation | Semantic chunks with overlap + synthesis |
| Visual structure | Wall of chat text | Three bullets | Concept trees, timeline rails, live stepper |
| Follow-up questions | Ungrounded chat | Usually none | Grounded in the current session summary and source |
| Local / private inference | Vendor account | Cloud-only | Ollama, LM Studio, or your endpoint |
| Architecture | Heavy web app | Monolithic content script | 30+ ES modules, zero npm, Manifest V3 |
Selectively tuned prompts and custom DOM parsers tailor each summary to the structural strengths of the medium.
Preserves real transcript timestamps, chapter markers, and narrative arc without hallucinating sequence.
Reads Chrome PDF text layers plus arXiv, PubMed, OpenReview, and IEEE pages with an academic research persona.
Employs Readability heuristics and a DOM accessibility fallback to strip ads, sidebars, and boilerplates.
Specialized DOM extractors for Coursera and Udemy extracting lesson objectives, lecture text, and code snippets.
Highest priority trigger: isolates exact user selection for analyzing complex paragraphs, formulas, or docs.
Standard structured digest with high-level thesis, key takeaways, and narrative timeline.
Dissects core premises, empirical evidence, unstated assumptions, and methodology limits.
Progressive conceptual ladder translating difficult jargon into intuitive analogies.
Identifies opposing tensions, counterarguments, and steelmanned viewpoints.
Extracts formal definitions, active recall flashcards, and retention takeaways.
Hierarchical tree with numbered points and nested sub-headings for quick scanning.
Vertical milestone rail with timestamp pills, phase markers, and event cards.
Filterable Core / Important / Supporting concept tree rendered from the heading contract.
Comprehensive overview synthesizing the entire source into 2-3 dense, clear paragraphs.
Actionable, bulleted insights formatted with bold anchors for high visual scannability.
Sequential walk-through of the argument with exact timestamps (e.g. [12:34]).
Deep-tier synthesis linking root causes, second-order effects, and architectural tradeoffs.
Concrete benchmarks, empirical experiments, and specific case studies mentioned.
Auto-generated questions that prompt the user to explore deeper nuance via chat.
Native multimodal and YouTube video support with high context windows and streaming chunk tokenization.
Connects to OpenAI GPT-4o, Anthropic via proxies, Groq, or OpenRouter with configurable base URLs.
Runs 100% locally with zero internet access required. Connects to Ollama, LM Studio, or vLLM endpoints.
The processing pipeline is partitioned into distinct, auditable stages. Each module enforces strict input/output contracts to eliminate silent failures.
Priority dispatcher: selection → PDF/paper → YouTube → course → Readability webpage.
Splits large transcripts at timestamp and paragraph boundaries with 1-sentence overlap.
Injects source metadata, depth guidelines, and heading contracts into a unified envelope.
Streams chunks via AbortController with per-tab request cancellation.
Evaluates section coverage scores; triggers targeted single-pass repair if shallow.
Every provider implements the normalized method signature generateText(prompt, settings, onChunk?). The provider has zero knowledge of tab IDs, user interface state, or section parsing. If a provider fails or the user cancels generation, an AbortController signal cleanly terminates in-flight fetch streams without leaking worker memory.
Unlike basic character chunking that cuts sentences mid-word, DeepDigest splits content based on structural hierarchy: timestamp segment boundaries → paragraph breaks → sentence delimiters → clause commas. Each chunk retains 1 sentence of rolling overlap to ensure continuity during final multi-chunk synthesis.
When generating in Deep or Long modes, the quality gate inspects the parsed AST for section depth, minimum bullet counts, and placeholder avoidance. If a specific section (e.g. Reasoning, Evidence & Claim Audit) falls below the threshold, a targeted repair prompt regenerates only that section — preserving healthy content and saving 80% of token compute.
The Chrome side panel keeps only the active session result and follow-up context in memory. Switching tabs clears the current panel state, avoiding persistent storage work and stale cross-tab data.
Explore how prompt templates are assembled dynamically. A shared outer envelope enforces system directives, language injection, and the strict markdown heading contract.
# SHARED ENVELOPE (lib/prompts/common.js)
You are an expert analytical research assistant. Your task is to transform the provided source content into an exceptionally well-structured, authoritative, and grounded synthesis.
## CORE DIRECTIVES:
1. Grounding: Rely strictly on the provided source content. Do not extrapolate unsupported claims.
2. Structure: Follow the mandatory section headings exactly as specified.
3. Language: Respond in the requested output language: {outputLanguage}.
4. Tone: Analytical, concise, objective, and dense with actionable insight.
## MANDATORY HEADING CONTRACT:
## Main Summary
## Executive Takeaways
## Details of the Video / Content
## Connections, Causes & Tradeoffs (for Deep depth)
## Reasoning, Evidence & Claim Audit
## Caveats, Biases & Open Questions
## Follow-up Questions
{sourceSpecificTemplate}
Every technical architectural decision was made to balance UX responsiveness, maintainability, and resource utilization.
Decision: Avoided React/Vite/Webpack; wrote the extension in pure ES2020+ modules.
Rationale: Chrome extensions with MV3 service workers and side panels run lighter without hydration overhead. Instant reload during development and zero dependency vulnerability churn.
Decision: Used markdown headings (## Executive Takeaways) instead of JSON Schema generation.
Rationale: Allows incremental streaming rendering to the user within 300ms. JSON requires buffering the full payload before parsing, which degrades perceived performance.
Decision: Repaired only failing sections (1 pass maximum) instead of re-running the prompt.
Rationale: Avoids discarding already well-generated sections and cuts token usage and user latency by 75% on edge-case summaries.
Decision: Centralized all defaults, bounds, and normalizations into settings-schema.js.
Rationale: Both the side panel and the options page consume the same schema, preventing migration bugs and guaranteeing forward compatibility.
Extraction runs only after an explicit action. Credentials stay in the service worker. Local providers never leave the machine.
Provider API keys never enter the content script or the page DOM. The background worker is the only process that reads them.
Point the local provider at 127.0.0.1. Internal docs and selections can be summarized without a cloud round-trip.
Nothing is extracted until you click Generate, use the context menu, or press Ctrl+Shift+S.
Results, follow-up chat, and workflow progress stay in the active side-panel session only. Switching tabs or closing the panel clears them; no hidden reading history.
Install with 1-click directly from the official Firefox Add-ons store, or load unpacked in Chrome developer mode in just a few minutes.
Published and verified on Mozilla Add-ons (AMO). One-click install with automatic updates.
Clone the repository or download the ZIP, then open the summarizer-extension folder.
git clone https://github.com/thaihai-swe/browser-extensions.git
Go to chrome://extensions and turn on Developer mode in the top-right corner.
Click Load unpacked and select the .output/chrome-mv3 directory.
Open a YouTube video, arXiv paper, PDF, or article, then press Ctrl+Shift+S (Cmd+Shift+S on macOS) or use the context menu.
Gemini, OpenAI, or a local endpoint needs configuration in Options. Local models run without a cloud key. Provider usage may incur charges from that provider, not from this project.
Short answers for setup, privacy, long videos, and what the extension does — and does not — do.
No. You can use a Gemini key from Google AI Studio, an OpenAI-compatible endpoint, or a fully local model through Ollama or LM Studio. There is no Chrome Web Store fee and no account on this project.
Typical pages use one provider request. Long transcripts can be split into up to four semantic chunks with sentence overlap, then synthesized. Deep and Long modes may run one targeted repair pass if a section is too shallow. A four-phase stepper shows Extract → Analyze → Synthesize → Quality as this happens.
Yes. The extractor detects Chrome PDF text layers, .pdf URLs, and paper hosts such as arXiv, PubMed, OpenReview, and IEEE. Academic sources use a research-scientist persona covering Research Question, Methodology, Empirical Results, and Caveats. Long papers can chunk up to 60,000 characters.
Yes. Selected text has priority over PDF, YouTube, course, and webpage extractors. Pair selection mode with a local provider if the content should never leave the machine.
Ctrl+Shift+S (or Cmd+Shift+S) opens the side panel and starts a summary. You can also click the toolbar icon or right-click Summarize with DeepDigest.
Named presets live in Options. They keep DeepDigest’s heading contract and grounding rules while adding your instructions. Placeholders such as __CONTENT__ and __LANG__ are supported.
No analytics and no hidden scrape. Content is sent to the provider you configured, and only after you generate. Results and follow-up context remain session-only and are cleared when the panel session ends.
The project includes 14 dedicated documentation guides covering every subsystem from content pipelines to test matrices.
DeepDigest showcases how product design, AI prompt engineering, and clean systems architecture combine to create a reliable cognitive companion for developers and researchers.