Beta Release — This is the first public beta. Expect rough edges, evolving APIs, and occasional AI generation failures. We welcome bug reports and feedback via GitHub Issues.
A feature-rich, open-source desktop application inspired by Google's NotebookLM. Built with Electron, React, and powered by multi-provider AI (Gemini, Claude, OpenAI, Groq). Upload documents, chat with your sources, and generate studio-quality content — AI video overviews & music videos (Veo 3.1), AI podcasts, HTML presentation decks with a structured slide editor, image slide decks with drag-and-drop editing, PPTX export, whitepapers, infographics with video animation, flashcards, quizzes, mind maps, dashboards, literature reviews, competitive analyses, reports, and more — with agentic RAG, multi-agent generation pipelines, voice Q&A, cross-session memory, Knowledge Hub with file connectors and knowledge graph, and local embeddings.
- Features Overview
- What's New — Beta
- Tech Stack
- Getting Started
- Application Guide
- AI Architecture
- Clipboard Quick-Capture
- Keyboard Shortcuts
- Project Structure
- Database Schema
- License
| Category | Features |
|---|---|
| Source Ingestion | PDF, DOCX, TXT, Markdown, Website URLs, YouTube transcripts, Audio files (MP3/WAV/M4A/OGG/FLAC), Paste text |
| AI Chat | Streaming responses, source-grounded citations, conversation history, suggested prompts, 6 interactive artifact types (Table, Chart, Mermaid, Kanban, KPI, Timeline), one-click artifact shortcut chips, voice Q&A |
| Deep Research | Multi-step AI analysis with real-time progress updates |
| Video Overview | AI-planned narrated explainer videos from notebook sources — scene planning, image generation, Veo animation (5 models), narration sync, Ken Burns fallback, ffmpeg assembly |
| Music Video | Upload MP3/WAV + optional lyrics — AI generates visuals timed to your audio with automatic duration detection |
| Audio Overview | Multi-speaker AI podcast with 4 format styles (Deep Dive, Brief, Critical Analysis, Debate) |
| HTML Presentations | Structured slide editor with slide navigator, layout picker (title, content, two-column, card-grid, stat-row, quote, closing), per-slide AI regeneration, Nano Banana image generation for slide blocks, live preview, PPTX export with embedded images, prompt templates |
| Image Slides | AI-generated slide decks with 6 visual styles + custom style builder, 3 image models (Nano Banana Pro / 2 / Classic), 2 render modes (full-image / hybrid editable), 5 aspect ratios (16:9, 4:3, 1:1, 3:4 portrait, 9:16 portrait), visual-first approach with text integrated into imagery, fullscreen presenter, rich text drag-and-drop editor, Save button with feedback, PDF export with text overlays, prompt customization, AI-recommended slide count |
| White Paper | AI-generated multi-section white papers with cover image, section illustrations, table of contents, references, 3 image models, and A4 PDF export |
| Infographic | AI-generated single-page visuals in 3 format types (Infographic, Advertisement, Social Post), full-image or hybrid mode, 5 aspect ratios including portrait (9:16, 3:4), custom color palettes, 3 image model options, and Veo 3.1 video animation |
| Study Tools | Flashcards, Quizzes (multiple choice), Reports, Mind Maps, Data Tables, Dashboards, Literature Reviews, Competitive Analyses, Document Comparisons, Citation Graphs — search/filter/sort for all generated content |
| Notes | Obsidian-class rich editor (Tiptap), slash command menu (/), callout blocks (5 types), toggle/collapsible sections, text color picker, folders, FTS5 search, outline panel, [[wiki links]] with backlinks, #tags, daily notes, templates, convert to source |
| Knowledge Graph | Force-directed visualization of all notes and their [[wiki link]] connections, tag-colored nodes, zoom/pan/drag |
| Knowledge Wiki | Karpathy LLM Wiki concept — AI-maintained entity/concept/topic pages from sources, coverage indicators, confidence scores, lint engine |
| Tasks | Auto-extracted from note checkboxes, vault-wide queries, due dates (due:YYYY-MM-DD), priority (!high/!medium/!low), two-way sync |
| Kanban Board | Tag-based columns, drag-and-drop cards, visual note organization |
| Canvas | Infinite spatial workspace powered by tldraw — shapes, arrows, freehand, text, auto-save |
| AI Note Features | Auto-tag suggestions, auto-link detection, note summarization (short/medium/long) |
| Workspace | Link local folders, file tree browser, text editor with AI rewrite, .gitignore support |
| Export | Download audio as WAV, slides as PNG, slide decks as PDF (with text overlays for hybrid mode), whitepapers as A4 PDF, copy reports to clipboard |
| AI Architecture | Agentic RAG (multi-query retrieval), multi-agent generation pipeline (Research → Write → Review), AI output validation with retry, cross-session memory |
| Embeddings | Tiered embedding system: local ONNX (all-MiniLM-L6-v2) → Gemini API → hash fallback |
| Knowledge Hub | File connectors for local folder ingestion, knowledge graph visualization, Ollama-powered local embeddings, knowledge search across all indexed content |
| Multi-Provider Chat | Gemini (default), Claude, OpenAI, Groq — switch providers and models in Settings |
| Productivity | Clipboard quick-capture via system tray (Cmd+Shift+N), chat-to-source pipeline, smart source recommendations |
This beta includes 20+ major features transforming DeepNote AI from a notebook tool into a full intelligent research platform:
| Feature | Description |
|---|---|
| 17 Studio Tools | Video Overview, Music Video, Audio, Image Slides, HTML Presentations, Flashcards, Quiz, Report, Mind Map, Data Table, Dashboard, Literature Review, Competitive Analysis, Document Comparison, Citation Graph, Infographic, White Paper, and Deep Research |
| Video Overview & Music Video | Two-phase pipeline (storyboard review → animation) with 5 Veo models (3.1/3.1 Fast/3/3 Fast/2), narration sync, mood/style control, resolution picker (720p–4K), quota-aware fallback, ffmpeg assembly |
| HTML Presentation System | Structured slide editor with 7 layout types, per-slide editing and AI regeneration, slide navigator with drag-and-drop reordering, Nano Banana image generation for individual slide blocks, live preview with instant refresh, and PPTX export with embedded images |
| Prompt Templates | Create, save, and manage custom prompt templates for slide deck generation — reuse your best prompts across notebooks |
| Slide Deck Enhancements | Visual-first approach (text integrated into imagery), 5 aspect ratios including portrait (3:4, 9:16), AI-recommended page count (up to 20 slides), prompt customization, improved generation reliability |
| Infographic Animation | Generate video animations from infographics using Google Veo 3.1 — image-to-video with 4-8s clips, up to 4K resolution |
| Knowledge Hub | Replaced DeepBrain with local-first Knowledge Hub — file connectors for folder ingestion, knowledge graph visualization, Ollama-powered local embeddings |
| White Paper Generator | Multi-section academic/business/technical papers with cover images, section illustrations, ToC, references, and A4 PDF export |
| Infographic Generator | 3 format types (Infographic, Advertisement, Social Post) with 5 aspect ratios including portrait (3:4, 9:16), visual-first design, full-image or hybrid mode, 6 style presets + custom style builder with reference image analysis |
| Hybrid Slide Editor | Drag-and-drop text overlays on AI backgrounds, Save button with visual feedback, arrow key navigation respects text editing, PDF export includes text overlays |
| Custom Style Builder | Upload a reference image — the AI extracts and replicates its visual style for slides, infographics, and whitepapers |
| Image Model Picker | Choose between 3 Gemini image models per creation: Nano Banana Pro (highest fidelity), Nano Banana 2 (fast & efficient), Nano Banana (speed optimized) |
| Glass Morphism UI | Polished dark/light theme with frosted glass effects, refined modals, and fullscreen dialogs |
| Agentic RAG | Multi-query retrieval: AI generates 2-3 targeted sub-queries, deduplicates results, checks sufficiency |
| Multi-Agent Pipeline | Complex content goes through Research → Write → Review pipeline with automatic revision if quality < 6/10 |
| AI Output Validation | Middleware validates all JSON output, strips markdown fences, retries with error feedback (3 attempts) |
| Cross-Session Memory | AI remembers preferences and learning patterns per notebook with confidence scoring |
| Voice Q&A | Audio transcription → RAG chat → TTS response — speak to your sources |
| Multi-Provider Chat | Gemini, Claude, OpenAI, Groq — switch providers and models per conversation |
| Local ONNX Embeddings | Offline embeddings via all-MiniLM-L6-v2 with tiered fallback (ONNX → Gemini → hash) |
| Global Search | System-wide search across all notebooks, sources, files, emails, and memories (Cmd+K) |
| Clipboard Quick-Capture | System tray icon + Cmd+Shift+N to capture clipboard content to your notebook |
| Layer | Technology |
|---|---|
| Framework | Electron 39 via electron-vite |
| Video | fluent-ffmpeg, @ffmpeg-installer/ffmpeg, @ffprobe-installer/ffprobe |
| Frontend | React 19, TypeScript 5, Tailwind CSS v4 |
| State | Zustand |
| Routing | react-router-dom v7 |
| Database | SQLite (better-sqlite3 v12.6) + Drizzle ORM |
| AI | Google Gemini, Anthropic Claude, OpenAI, Groq (@google/genai, multi-provider) |
| Local Embeddings | ONNX Runtime (all-MiniLM-L6-v2, 384-dim) |
| Rich Text | Tiptap (ProseMirror-based) |
| Drag & Drop | react-draggable |
| Document Parsing | pdf-parse, mammoth (DOCX) |
| Icons | lucide-react |
| Font | Inter (bundled via @fontsource-variable) |
AI Models Used:
gemini-3-flash-preview(default) /gemini-3-pro-preview/gemini-2.5-flash/gemini-2.5-pro— Chat, content generation, document analysis, agentic RAGclaude-sonnet-4-6/claude-opus-4-6/claude-haiku-4-5— Claude chat providersgpt-4o/gpt-4o-mini— OpenAI chat providersllama-3.3-70b-versatile— Groq chat providergemini-2.5-flash-preview-tts— Multi-speaker text-to-speech (Kore & Puck voices)gemini-3-pro-image-preview(Nano Banana Pro) /gemini-3.1-flash-image-preview(Nano Banana 2) /gemini-2.5-flash-image(Nano Banana) — AI image generation for slides, infographics, whitepapers — selectable per creationveo-3.1-fast-generate-preview(default) /veo-3.1-generate-preview/veo-3-fast-generate/veo-3-generate/veo-2-generate— Video generation (image-to-video animation, 4-8s clips, up to 4K)text-embedding-004— Text embeddings for RAG (Gemini cloud tier)all-MiniLM-L6-v2— Local ONNX embeddings (offline tier, 384-dim)
- Download the latest DMG from Releases
- Open the DMG and drag DeepNote AI to your Applications folder
- Launch the app, open Settings (gear icon), and enter your Gemini API Key
- Click Test to verify, then Save — you're ready to go
Note: The app is signed but not notarized. On first launch macOS may show a security warning — right-click the app and choose "Open" to bypass it.
- Node.js 20+ and npm
- Google Gemini API Key (free at aistudio.google.com)
# Clone the repository
git clone https://github.com/Clemens865/DeepNote-AI.git
cd DeepNote-AI
# Install dependencies (ignore scripts to avoid premature native builds)
npm install --ignore-scripts
# Rebuild native modules for Electron
npx @electron/rebuild -f -w better-sqlite3- Launch the app
- Click the Settings icon (gear) in the header
- Enter your Gemini API Key
- Click Test to verify, then Save
The API key is stored locally and never leaves your machine (except to call the Gemini API).
# Development mode with hot reload
npm run dev
# Preview production build
npm run start# Full build (typecheck + bundle)
npm run build
# Platform-specific distributables
npm run build:mac # macOS .dmg
npm run build:win # Windows .exe
npm run build:linux # Linux .AppImageThe Dashboard is the home screen. Each notebook is an isolated workspace with its own sources, chat history, notes, and generated content.
- Create a notebook by clicking the "+" button - choose a title and icon
- Workspace notebooks can be linked to a local folder for file editing
- Search notebooks by title using the search bar
- Delete notebooks via the context menu
Each notebook displays the number of sources and the last-modified date.
Sources are the knowledge base for your notebook. All AI features (chat, studio) draw from selected sources.
Supported Source Types:
| Type | Description |
|---|---|
| File Upload | PDF, DOCX, DOC, TXT, MD files |
| Website | Extracts text content from any URL |
| YouTube | Extracts transcript from YouTube video URLs via Gemini AI |
| Audio | Transcribes MP3, WAV, M4A, OGG, FLAC, AAC files via Gemini multimodal |
| Paste Text | Directly paste text content with a custom title |
Source Features:
- Select / Deselect - Toggle which sources provide AI context (checkbox per source)
- Source Guide - AI-generated summary overview of each source
- Auto-chunking - Documents are split into chunks with embeddings for RAG retrieval
- Delete sources individually
When viewing a source, DeepNote AI can recommend related sources from your other notebooks using vector similarity search. This helps you discover connections between different areas of research.
The Chat Panel is the main interaction mode. Ask questions and get AI-generated answers grounded in your selected sources.
- Streaming responses - Answers appear token-by-token in real time
- Source citations - Responses include
[Source N]references linking to specific source chunks - Suggested prompts - Quick-start buttons: "Summarize my sources", "Key takeaways", "Create a study guide"
- Upload from chat - Add new documents or audio files directly from the chat input
- Save to Note - Save any AI response as a note for later reference
- Clear history - Reset the conversation
- Cross-session memory - The AI remembers your preferences and learning patterns
- Knowledge Hub integration - When Knowledge Hub is configured with file connectors, the AI can search across all ingested knowledge. Results appear as context-enriched responses grounded in your knowledge base
Interactive Artifacts:
The AI can embed rich visualizations directly in its responses:
| Artifact | Description |
|---|---|
| Table | Sortable data tables with columns and rows |
| Chart | Bar, line, or pie charts with tooltips and legend |
| Mermaid Diagram | Flowcharts, sequence diagrams, ER diagrams |
| Kanban Board | Task cards with assignee, priority, and status |
| KPI Cards | Metric cards with progress bars and sentiment colors |
| Timeline | Horizontal timeline with dated events |
Artifact Shortcut Chips:
Six quick-action buttons appear above the chat input (when sources are selected). Click any chip to instantly request that artifact type - Table, Chart, Diagram, Kanban, KPIs, or Timeline.
Click the microphone button next to the chat input to start a voice conversation with your sources.
- Record - Tap the mic button to start recording your question
- Transcription - Audio is transcribed via Gemini AI
- RAG-powered response - Your question is answered using selected source context
- TTS playback - The AI's response is spoken back to you
- Transcript display - Full text transcript of the conversation appears in the overlay
- Multiple turns - Continue asking follow-up questions in the same session
Every AI response has action buttons to integrate it into your workflow:
| Action | Description |
|---|---|
| Note | Save the response as a note in your notebook |
| Source | Add the response as a new source (with full embedding ingestion) |
| Workspace | Write the response as a Markdown file to your workspace folder |
| Generate | Send the response to studio as context for content generation |
| Copy | Copy the response text to clipboard |
Deep Research performs a multi-step, in-depth analysis of your sources. It goes beyond simple chat by running a longer AI pipeline with real-time progress updates.
- Click the Deep Research button in the chat panel
- Enter a research query
- Watch progress indicators as the AI analyzes sources in depth
- Results appear as a comprehensive research report in the chat
The Notes Panel provides a scratchpad for your own thoughts alongside AI-generated content.
- Create notes with the "+" button
- Auto-save - Content saves automatically as you type (500ms debounce)
- Editable titles - Click any note title to rename it
- Convert to Source - Turn a note into a source document so the AI can reference it in chat and studio
- Timestamps - Each note shows when it was last edited
- Delete notes with confirmation
The Studio Panel contains 17 AI-powered content generation tools. Each tool transforms your selected sources into a different output format.
All studio tools support:
- Custom instructions - Guide the AI's focus and audience
- Length options - Short, Default, or Long output
- Rename generated content
- Delete generated content
- Generation history - All past outputs are accessible
- Search & filter - Find generated items by title, filter by type, sort by newest/oldest/title/type
- Multi-agent pipeline - Complex types (report, whitepaper, literature review) use a Research → Write → Review pipeline with real-time progress indicators
Generate AI-powered narrated videos or music videos from your notebook sources.
Two Modes:
| Mode | Description |
|---|---|
| Video Overview | AI plans scenes from your sources, generates narrated explainer video with images, animation, and voiceover |
| Music Video | Upload an MP3/WAV file (+ optional lyrics), AI generates visuals timed to the music duration |
Two-Phase Pipeline:
- Storyboard (Phase 1) — AI plans scenes → generates images (4 parallel) → presents storyboard for review
- Animation (Phase 2) — After approval, generates narration (TTS) → animates scenes (Veo) → assembles final MP4 (ffmpeg)
Storyboard Review:
- Visual grid of all scene thumbnails
- Per-scene regeneration with custom instructions
- Approve and proceed to animation, or regenerate individual scenes
5 Veo Video Models:
| Model | Price/sec | Resolution | Audio | Duration |
|---|---|---|---|---|
| Veo 3.1 Fast | $0.15 | 720p/1080p/4K | Yes | 4-8s |
| Veo 3.1 | $0.40 | 720p/1080p/4K | Yes | 4-8s |
| Veo 3 Fast | $0.15 | 720p/1080p | Yes | 8s |
| Veo 3 | $0.40 | 720p/1080p | Yes | 8s |
| Veo 2 | $0.35 | 720p | Silent | 5-8s |
Video Features:
- Narrative styles — Explain, Present, Storytell, Documentary
- Narration sync — TTS audio fitted to match each video clip duration
- Mood control — Auto, custom prompt, or reference image style
- Style presets — 6 visual presets + custom style builder with color picker
- Quota-aware — Auto-fallback between Veo model variants on 429; Ken Burns fallback on full exhaustion
- Music video auto-detection — Detects audio file duration via ffprobe, generates exact scene count
- Resolution selector — 720p, 1080p, 4K (filtered by model capability)
- Progress tracking — Per-scene progress with retry status and fallback indicators
Generate a podcast-style audio conversation between two AI speakers discussing your source material.
Format Options:
| Format | Description |
|---|---|
| Deep Dive | Comprehensive discussion covering all key topics |
| Brief Overview | Short, high-level summary |
| Critical Analysis | Examines strengths, weaknesses, and implications |
| Debate | Two speakers take opposing perspectives |
Length Options:
- Short: 8-12 conversational turns
- Default: 15-25 turns
- Long: 30-45 turns
Audio Features:
- Two distinct AI voices (Kore & Puck) via Gemini TTS
- Built-in audio player with play/pause/seek
- Full transcript view with color-coded speakers
- Download as WAV file
The most advanced studio feature. Generates complete slide presentations with AI-created background images and editable text overlays.
6 Visual Style Presets:
| Style | Background | Accents | Aesthetic |
|---|---|---|---|
| Blueprint Dark | Dark teal | Orange & Cyan | Sci-fi command center, circuit diagrams |
| Editorial Clean | Cream | Orange & Blue | Magazine editorial, hand-drawn sketches |
| Corporate Blue | White | Navy & Light Blue | Boardroom professional, geometric icons |
| Bold Minimal | White | Black & Red | Apple keynote, large typography |
| Nature Warm | Cream | Green & Terracotta | Botanical, organic patterns |
| Dark Luxe | Black | Gold & White | Premium luxury, elegant lines |
Custom Style - Upload a reference image and the AI will extract and replicate its visual style.
Image Model Selection:
Each generation lets you choose between three Gemini image models:
| Model | Gemini ID | Best For |
|---|---|---|
| Nano Banana Pro | gemini-3-pro-image-preview |
Highest fidelity, advanced reasoning for complex prompts |
| Nano Banana 2 | gemini-3.1-flash-image-preview |
Fast & efficient, great for high-volume use |
| Nano Banana | gemini-2.5-flash-image |
Speed optimized, lowest latency |
Presentation Formats:
- Presentation - Educational / informational slides
- Pitch Deck - Business pitch (Problem, Solution, Market, Ask)
- Report Deck - Data-driven analysis with charts
Slide Count: Test (3), Short (5), Default (10), Recommended (AI-optimized, up to 20)
Aspect Ratios: 16:9 (widescreen), 4:3 (classic), 1:1 (square), 3:4 (portrait), 9:16 (tall portrait)
Two Render Modes:
| Mode | Description |
|---|---|
| Full Image | Visual-first — text integrated into the imagery as part of the composition |
| Hybrid (Editable) | AI background image + draggable HTML text overlays |
Rich Slide Editor (Hybrid Mode):
Click the pencil icon on any hybrid slide to enter edit mode:
- Drag & Drop - Freely position title, bullet, and text elements anywhere on the slide
- Rich Text Formatting - Bold, italic, underline, strikethrough via Tiptap editor
- Bullet Lists - Toggle bullet list formatting
- Text Alignment - Left, center, right alignment per element
- Font Size - Dropdown with sizes from 12px to 48px
- Text Color - 10-color palette picker
- Links - Add hyperlinks to text
- Add Text Box - Create new free-positioned text elements
- Delete Elements - Remove unwanted text boxes (minimum 1 must remain)
- Resize - Drag the corner handle to adjust element width
- Auto-save - All changes persist automatically (800ms debounce)
- Save Button - Explicit save with visual feedback (saving spinner → "Saved" checkmark)
- Click outside to deselect elements
Presentation Features:
- Fullscreen Mode - Press the expand button for a full-screen presentation
- Keyboard Navigation - Arrow keys and Space to advance slides (disabled while editing text to preserve cursor position)
- Thumbnail Strip - Visual slide navigator
- Speaker Notes - Show/hide AI-generated speaker notes per slide
- Download - Save individual slides as PNG files
Generate structured HTML presentations with a full slide editor. This is the newest presentation format, offering fine-grained per-slide editing with AI assistance.
7 Slide Layouts:
| Layout | Description |
|---|---|
| Title Slide | Centered title + subtitle with accent bar |
| Section Header | Section divider with underline accent |
| Content | Title + body content (text, bullets, images) |
| Two Column | Side-by-side content layout |
| Card Grid | 2-3 column card grid (up to 6 cards) |
| Stat Row | Key metrics displayed as large numbers |
| Quote | Centered quote with attribution |
| Closing | Summary slide with checkmark bullets |
Slide Editor Features:
- Preview / Edit toggle - Switch between live HTML preview and per-slide editor
- Slide Navigator - Thumbnail sidebar with drag-and-drop reordering, add/delete slides
- Per-slide editing - Edit title, subtitle, speaker notes, and individual content blocks
- Layout picker - Change slide layout type on the fly
- AI Regeneration - Regenerate any slide with custom instructions
- Image Generation - Add image blocks and generate AI images via Nano Banana models
- Live Preview - Instant HTML preview refresh on every edit
- PPTX Export - Export to PowerPoint with embedded images, themed colors, and template support
- Prompt Templates - Save and reuse custom generation prompts
Generate AI-powered single-page visuals from your sources with format-specific design intelligence.
3 Format Types:
| Format | Description |
|---|---|
| Infographic | Data visualization with icons, charts, key metrics, and flow diagrams — like a conference poster or editorial dashboard |
| Advertisement | Bold marketing asset with dominant headline, striking visual metaphor, and call-to-action — like a magazine ad or product launch hero |
| Social Post | Shareable visual with hero image and single bold statement — optimized for Instagram, LinkedIn, or Twitter |
Aspect Ratios: 16:9, 4:3, 1:1, 3:4 (portrait), 9:16 (tall portrait — ideal for social stories)
Render Modes:
- Full Image - Visual-first: all text and graphics integrated into a single AI-generated image
- Hybrid - AI background with structured HTML text annotations overlaid
Style Options: 6 visual presets + custom style builder (upload a reference image for AI style matching)
Additional Features:
- Image Model Picker - Choose between Nano Banana Pro, Nano Banana 2, or Nano Banana
- Veo Animation - Generate 4-8s video clips from infographics via Google Veo 3.1
- Custom Color Palettes - Extract or define color schemes for consistent branding
- Download - Save as PNG image
Generate multi-section academic, business, or technical papers with AI-generated illustrations.
- Cover image - AI-generated cover art matching the paper's theme
- Section illustrations - Each major section includes a generated image
- Table of Contents - Auto-generated with clickable navigation
- References - AI-generated reference list based on source material
- 3 image models - Choose Nano Banana Pro, 2, or Classic per generation
- A4 PDF Export - Professional document layout with images, typography, and formatting
Generate interactive flashcards from your sources for study and review.
- Card Count: Fewer (5-8), Standard (10-20), More (25-35)
- Difficulty: Easy, Medium, Hard
- Interactive - Click cards to flip and reveal the answer
- Source References - Each card cites the relevant source
Generate multiple-choice quizzes to test your understanding.
- Question Count: Fewer (3-5), Standard (5-10), More (12-20)
- Difficulty: Easy, Medium, Hard
- 4 options per question
- Interactive answer reveal - Green for correct, red for incorrect
- Explanations - Each answer includes a "why this is correct" explanation
Generate a structured written report from your sources.
- Executive Summary + 3-6 detailed sections
- Custom focus - Use instructions to target specific topics or audiences
- Multi-agent pipeline - Research → Write → Review for higher quality
- Copy to clipboard - One-click copy of the full report text
Generate a hierarchical mind map visualization of your source content.
- 3 levels deep - Central topic, main branches, sub-branches
- 3-6 main branches with 2-4 children each
- Color-coded depth - Indigo (root), Blue (L1), Emerald (L2), Amber (L3)
- Expandable tree - Click to expand/collapse branches
Extract structured data from your sources into a sortable table.
- 5-15 rows of meaningful data points
- AI-optimized columns - The AI determines the best column structure
- HTML table display with clean formatting
Link a local folder to your notebook and browse, edit, and AI-index files without leaving the app.
File Tree:
- Hierarchical browser - Navigate your project's file structure
- .gitignore support - Respects
.gitignorepatterns for file exclusion - Status indicators:
- Gray - Unindexed (not included in AI context)
- Green - Indexed (included in AI context)
- Yellow - Stale (file changed on disk since last index)
- Red - Error during indexing
- Select / Deselect files - Control which files become AI sources
File Editor:
- Text editing - Edit any text file directly in the app
- Dirty state indicator - Orange dot when there are unsaved changes
- Auto-detect read-only - Binary/unsupported files open in read-only mode
AI Rewrite:
- Select text in the editor and press
Cmd/Ctrl + K - An inline popup appears near your selection
- Type an instruction (e.g., "Make it more concise", "Add error handling")
- The AI rewrites just the selected text in context of the full file
- One-click replace of the original selection
Workspace Sync:
- Detects file changes (added, modified, deleted) on disk
- Re-indexes stale files to keep AI context up to date
- Shows sync status banner with counts
DeepNote AI uses several advanced AI patterns under the hood to deliver high-quality results.
Traditional RAG uses a single query to search for relevant context. DeepNote AI's Agentic RAG system reasons about your question first:
- Sub-query generation - The AI generates 2-3 targeted search queries from your question
- Multi-query search - Each sub-query is embedded and searched independently
- Deduplication & ranking - Results are merged by chunk ID and ranked by combined score
- Sufficiency check - The AI evaluates whether enough context was retrieved, and can request one additional retrieval round
This produces significantly better context for complex, multi-faceted questions.
For complex studio content types (report, literature review, competitive analysis, dashboard, whitepaper), generation uses a three-stage pipeline:
| Stage | Role | Description |
|---|---|---|
| Research | Analyst | Extracts key themes, facts, data points, and relationships from source material |
| Write | Writer | Generates structured content using research findings + source material |
| Review | Reviewer | Validates quality, checks structure, grades output (1-10). If score < 6, feeds back to writer for revision |
Real-time progress updates show which stage is currently executing.
All AI-generated structured content passes through a validation middleware pipeline:
- JSON validation - Strips markdown fences, validates JSON parsing
- Structure validation - Per-type schema checks (quiz must have questions with options, flashcards must have front/back, etc.)
- Retry with feedback - On validation failure, the error message is appended to the prompt and the AI retries (up to 3 attempts)
The AI learns from your conversations and remembers preferences across sessions:
- Automatic extraction - After each conversation, the AI extracts preferences, learning patterns, and context
- Memory types - Preference, learning, context, and feedback memories
- Confidence scoring - Each memory has a confidence score (0.0-1.0) that updates over time
- Per-notebook + global - Memories can be scoped to a specific notebook or applied globally
- System prompt injection - Relevant memories are automatically included in the AI's context
DeepNote AI supports a tiered embedding system for flexibility between online and offline use:
| Tier | Model | Dimensions | Requirement |
|---|---|---|---|
| 1. Local ONNX | all-MiniLM-L6-v2 | 384 | Auto-downloads model (~80MB) |
| 2. Gemini API | gemini-embedding-exp-03-07 | 768 | API key + internet |
| 3. Hash fallback | Content hash | 768 | Always available |
The system automatically selects the best available tier. Configurable via Settings (auto, gemini, or local).
A system tray icon provides quick access to capture clipboard content without opening the main window.
- Global shortcut - Press
Cmd+Shift+Nfrom any app to capture clipboard text - Clipboard history - Last 10 captured items stored in the tray menu
- Add to notebook - Send any captured text directly as a source to your current notebook
- Always accessible - Works even when the main window is minimized
| Shortcut | Action |
|---|---|
Cmd/Ctrl + S |
Save file in workspace editor |
Cmd/Ctrl + K |
Open AI rewrite popup (with text selected) / Open global search |
Cmd + Shift + F |
Open global search |
Cmd + Shift + N |
Capture clipboard to notebook (global) |
Arrow Right / Space |
Next slide (fullscreen) |
Arrow Left |
Previous slide (fullscreen) |
Escape |
Close fullscreen / close modals / close voice overlay |
Enter |
Submit in dialogs |
src/
shared/
types/ # Shared TypeScript interfaces & IPC channel definitions
index.ts # Core types (Notebook, Source, Note, ChatMessage, UserMemory, etc.)
ipc.ts # IPC channel names & handler type map
main/ # Electron main process
index.ts # App entry, window creation, protocol registration, tray init
db/
index.ts # SQLite database initialization (auto-create tables)
schema.ts # Drizzle ORM schema definitions (including user_memory)
ipc/
chat.ts # Chat message handlers (streaming, agentic RAG, DeepBrain context)
knowledge.ts # Knowledge Hub handlers (file connectors, knowledge graph, search)
clipboard.ts # Clipboard history & add-to-notebook handlers
config.ts # API key management
memory.ts # Cross-session memory CRUD handlers
notebooks.ts # CRUD for notebooks
notes.ts # CRUD for notes
research.ts # Deep research pipeline
sources.ts # Source ingestion, parsing, embedding, recommendations
studio.ts # Content generation & image slides (with pipeline progress)
voice.ts # Voice Q&A session handlers
workspace.ts # File tree, editing, syncing
services/
ai.ts # Gemini AI (chat, generation, planning, memory-aware prompts)
agenticRag.ts # Multi-query agentic RAG system
aiMiddleware.ts # AI output validation & retry middleware pipeline
chunker.ts # Document chunking with token counts
config.ts # Persistent config store (API key, embeddings model)
documentParser.ts # PDF, DOCX, TXT/MD parsing
embeddings.ts # Gemini embedding generation
generationPipeline.ts # Multi-agent Research → Write → Review pipeline
htmlRenderer.ts # HTML presentation rendering with image embedding
imagen.ts # AI image generation for slides
knowledgeIngestion.ts # Knowledge ingestion pipeline
knowledgeStore.ts # Knowledge storage and retrieval
localEmbeddings.ts # ONNX local embedding inference
memory.ts # Cross-session memory service
rag.ts # Retrieval-Augmented Generation (with agentic option)
recommendations.ts # Cross-notebook source recommendations
sourceIngestion.ts # Shared source processing pipeline
tieredEmbeddings.ts # Tiered ONNX → Gemini → hash embedding system
tray.ts # System tray icon & clipboard quick-capture
pptxRenderer.ts # PPTX generation with embedded images
pptxTemplateParser.ts # PPTX template parsing for theme extraction
tokenTracker.ts # API cost tracking (LLM tokens, Veo per-second, image gen)
tts.ts # Multi-speaker text-to-speech
ffmpegAssembler.ts # Video assembly, Ken Burns fallback, narration sync
veo.ts # Veo video generation (5 models, quota-aware fallback)
videoOverview.ts # Video Overview pipeline orchestrator (storyboard + animation)
vectorStore.ts # In-memory vector similarity search
voiceSession.ts # Voice Q&A session management
webScraper.ts # URL content extraction
workspace.ts # Filesystem scanning & .gitignore
preload/
index.ts # Context bridge API (typed IPC methods)
renderer/
src/
App.tsx # Router & top-level routes
main.tsx # React entry point
styles/
globals.css # Tailwind v4 config, theme tokens, Tiptap styles
stores/
appStore.ts # Global app state (theme, settings modal)
notebookStore.ts# Active notebook state
workspaceStore.ts# Workspace file tree & editor state
hooks/
useNotebooks.ts # Notebook listing hook
components/
chat/ # ChatPanel, ChatInput, ChatMessage, ChatKnowledgeResults, VoiceOverlay
common/ # Button, Modal, Spinner, Toast, SettingsModal
dashboard/ # Dashboard, NotebookCard
layout/ # AppLayout, Header, ResizablePanel
notes/ # NotesPanel, NoteEditor
sources/ # SourcesPanel, SourceList, AddSourceModal
knowledge/ # KnowledgeHub, ConnectorsTab, KnowledgeGraphTab
studio/ # StudioPanel, ToolGrid, GeneratedContentView,
# ImageSlidesView, ImageSlidesWizard,
# HtmlPresentationView, PresentationSlideEditor,
# PresentationSlideNav, HtmlPresentationWizard,
# DraggableTextElement, SlideEditorToolbar,
# StudioCustomizeDialog
workspace/ # WorkspaceLayout, FileTreeView, FileTreeNode,
# FileEditor, WorkspaceSyncBanner
DeepNote AI uses SQLite with 10+ tables:
| Table | Description |
|---|---|
notebooks |
Notebook metadata (title, emoji, workspace path) |
sources |
Ingested documents (content, type, selection state, source guide) |
chunks |
Chunked text with token counts for RAG retrieval |
notes |
User-created notes (title, content, converted-to-source flag) |
chat_messages |
Chat history with citations and metadata (role, content, citations JSON, metadata JSON) |
generated_content |
Studio outputs (type, data JSON, status, source IDs) |
workspace_files |
Workspace file manifest (path, hash, status, linked source ID) |
user_memory |
Cross-session AI memory (type, key, value, confidence score, timestamps) |
knowledge_documents |
Knowledge Hub ingested documents with embeddings |
knowledge_connectors |
File connector configurations for knowledge ingestion |
All data is stored locally in ~/.config/deepnote-ai/ (or platform equivalent). No data is sent to external servers except Gemini API calls for AI features.
This is a beta release. The following issues are known and tracked for future releases:
| Category | Issue | Severity |
|---|---|---|
| Performance | Chat message list not virtualized — may slow with very long histories | Medium |
| Performance | Vector search uses linear scan (no indexing) — scales poorly with many sources | Medium |
| Caching | Audio cache and slide image cache grow unbounded — no automatic cleanup | Low |
| Caching | Config file re-read from disk on every access — no in-memory cache | Low |
| Error Handling | Some API errors silently swallowed via .catch(() => {}) — user sees no feedback |
Medium |
| Error Handling | ||
| Security | API keys stored as plaintext JSON in config file — should use OS keychain | Medium |
| UX | Modal dialogs lack focus trapping and ARIA labels | Low |
| UX | Toast notifications can stack/overlap | Low |
| Accessibility | Icon-only buttons missing aria-label in several components |
Low |
We welcome bug reports at GitHub Issues.
This project is provided as-is for educational and personal use.