Skip to content

fix: remove docs - #9

Merged
KylinMountain merged 47 commits into
mainfrom
bugfix/compile
Apr 11, 2026
Merged

fix: remove docs#9
KylinMountain merged 47 commits into
mainfrom
bugfix/compile

Conversation

@rejojer

@rejojer rejojer commented Apr 9, 2026

Copy link
Copy Markdown
Member

Summary

  • Unified summary frontmatter to doc_type + full_text, removed redundant fields (sources, brief, source_doc, doc_id)
  • Concept pages now link to summaries instead of raw PDF filenames
  • Fixed duplicate frontmatter bug in both summary and concept pages (prompt fix + strip fallback)
  • Improved query agent: use full_text field, restrict get_page_content to pageindex docs, add self-talk, concise answers
  • Fixed image path mismatch in pageindex JSON content
  • Removed page marker comments from short doc source markdown
  • Fixed warning suppression (markitdown overrides filters at import time)
  • Improved init prompts with explicit defaults, American English spelling
  • Various output formatting fixes (tool call spacing, step name colons, unicode ellipsis)

Test plan

  • openkb init shows correct prompts with defaults
  • openkb add short doc: clean single frontmatter, no page markers in source
  • openkb add long doc: correct image paths in JSON content
  • openkb query on short doc: reads source via read_file, no get_page_content
  • openkb query on long doc: uses get_page_content with targeted page ranges
  • No PyPDF2 deprecation warning during any operation
KylinMountain and others added 30 commits April 9, 2026 00:21
Add _CONCEPTS_PLAN_USER (create/update/related JSON structure) and
_CONCEPT_UPDATE_USER templates; add TestParseConceptsPlan tests.
- Restore markitdown[all] extras for docx/pptx/xlsx support
- Sanitize concept names to prevent path traversal in compiler
- Add path traversal guard in copy_relative_images
- Fix _write_concept duplicate append when frontmatter lacks sources key
- Remove dead write_wiki_files function
- Fix watcher thread race in _schedule_flush
- Warn when unimplemented --fix flag is used in lint command
- Harden CI publish workflow with environment gate and SHA-pinned actions
- Fix test_indexer to actually assert IndexConfig flag values
- Fix test_converter to test correct PDF code path (pymupdf, not markitdown)
- Use str.find() instead of str.index() in frontmatter parsing to avoid ValueError
- Add _backlink_summary: ensures summary pages link to all related concepts
- Add _backlink_concepts: ensures concept pages link back to source summaries
- _update_index auto-creates index.md if missing
- Both merge into existing sections instead of duplicating
Adds parse_pages() to expand page specs like "1-3,7" into sorted
deduplicated int lists, and get_page_content() to read per-page JSON
(sources/{doc}.json) and format output with optional image paths.
Includes path-traversal guard consistent with existing tools.
Replace _SUMMARY_USER, _CONCEPT_PAGE_USER, and _CONCEPT_UPDATE_USER to
request a JSON object with "brief" (one-line summary) and "content" (full
Markdown). Add TestParseBriefContent to tests/test_compiler.py.
Replace markdown source generation with per-page JSON from PageIndex
get_page_content; remove render_source_md, _render_nodes_source,
_relocate_images, and _IMG_REF_RE. Image relocation is now done inline
per page. Update tests to assert .json output and mock get_page_content.
…or all docs

Remove _pageindex_retrieve_impl and the pageindex_retrieve tool; add
get_page_content_tool that uses the local JSON-based page store for all
long documents. Update instructions and schema description accordingly.
… indexer

- Default model changed from gpt-5.4 to gpt-5.4-mini
- Indexer get_page_content no longer uses hardcoded 9999 fallback
- Infers page_count from structure end_index when doc lacks page_count field
- Added debug logging for doc keys and page_count diagnosis
…e backlink for short docs

- index.md entries now show (short) or (pageindex) type marker
- Query agent prompt updated: guides agent to read sources for detail
- Removed list_files tool from query agent (index.md is sufficient)
- Short doc summaries now have source_doc frontmatter linking to sources/
- Reverted list_wiki_files to only list .md files
- Fixed tests for model name change and agent tool count
Replace sources/brief/source_doc/doc_id/source fields with two
consistent fields: doc_type (short|pageindex) and full_text pointing
to the actual source content under sources/.
Concept pages now reference summaries/{doc}.md instead of raw PDF
filenames. Also strips frontmatter from LLM content during concept
updates to prevent duplicate YAML blocks. Removes unused
_find_source_filename.
Add hint that summaries may omit details. Update search strategy to
reference the full_text frontmatter field instead of hardcoded paths.
rejojer and others added 14 commits April 10, 2026 04:35
Remove frontmatter format from schema to avoid LLM copying it.
Add strip as fallback in _write_summary and _write_concept create path.
Replace PageIndex get_page_content with pymupdf-based convert_pdf_to_pages
for long doc JSON generation. All image paths now use sources/images/ prefix
relative to wiki root. Removes dependency on PageIndex for source content.
Query agent can now view images referenced in source documents via
get_image tool, which returns ToolOutputImage for the LLM to inspect.
Prompt updated to use images when questions involve figures or visuals.
Accept all origin/dev changes including: image support in query agent,
robust JSON parsing with json_repair, unicode concept name support,
section-based index operations, cloud/local page extraction fallback.
@KylinMountain
KylinMountain changed the base branch from dev to main April 11, 2026 02:53
@KylinMountain KylinMountain changed the title fix: compile pipeline, query agent, and frontmatter improvements Apr 11, 2026
@KylinMountain
KylinMountain merged commit 726336a into main Apr 11, 2026
@rejojer
rejojer deleted the bugfix/compile branch April 11, 2026 17:45
KylinMountain added a commit that referenced this pull request Jul 21, 2026
…es (#198)

* feat(api): delete a knowledge base (POST /api/v1/kb/delete + openkb delete-kb)

Physical, irreversible KB deletion with a type-the-name confirmation.

- config.delete_kb: rmtree the KB directory + unregister it from the global
  registry (known_kbs / kb_aliases / default_kb, via the new unregister_kb).
  Guards against deleting a non-KB path; tolerates a ghost registry entry
  (directory already gone) by just unregistering.
- POST /api/v1/kb/delete: confirm_name must equal kb, re-checked server-side.
  Lives in a new api_kbs_router (sibling of the config/graph/output routers) so
  api.py stays under the 800-line gate; delete_kb self-locks (config lock), no
  create_app closure needed.
- openkb delete-kb NAME: type-the-name prompt (or --yes) then delete.

Tests: endpoint (success + unregister, confirm mismatch, non-KB target) in
test_api.py; config (physical delete + unregister, refuse non-KB, ghost
tolerance) in test_config.py.

Claude-Session: https://claude.ai/code/session_01XMxbhmAkxxVV8CFWCZDBaG

* feat(api): delete a wiki page (POST /api/v1/page/delete) with backlink impact

Delete a concept/entity page and degrade references cleanly instead of leaving
them dangling:

- page_ops.delete_wiki_page: dry_run reports the backlink pages (whose inbound
  [[wikilinks]] would be demoted) without touching anything; execute removes the
  page under the KB ingest lock, strips its index.md entry outright
  (compiler.remove_doc_from_index — the entry is a [[link]] line), and demotes
  the now-dangling inbound [[links]] to plain text in ONLY those backlink pages
  (lint.fix_broken_links restrict_to). The page ref is "section/stem", validated
  to concepts/ or entities/ only (path-traversal-safe).
- New api_pages_router hosts POST /api/v1/page (read — relocated here from
  api.py to keep it under the 800-line gate) + POST /api/v1/page/delete.

Tests: dry-run impact + execute (page gone, inbound link demoted to text, index
entry removed, sibling kept), 404 on missing page, and rejection of unsafe /
non-editable refs.

Claude-Session: https://claude.ai/code/session_01XMxbhmAkxxVV8CFWCZDBaG

* feat(api): edit a wiki page (PUT /api/v1/page) + link context (POST /api/v1/page/links)

- page_ops.edit_wiki_page: replace a concept/entity page's BODY while preserving
  its code-managed OKF frontmatter (type/description/sources) verbatim; any
  frontmatter in the submitted content is dropped. Dead [[wikilinks]] the user
  typed are demoted to plain text on save (OpenKB keeps the wiki broken-link-free)
  and returned as ghosts_stripped so the UI can warn. Atomic write under the KB
  ingest lock.
- page_ops.page_link_context (POST /api/v1/page/links): outbound + inbound links
  for the edit-impact panel. Editing the body does not break either (links are
  path-based), so this is context, not a blocker — the honest impact model.
- PUT /api/v1/page added to api_pages_router.

Tests: edit preserves frontmatter + keeps a resolvable link + demotes a dead one
(reports ghosts_stripped) + replaces the body; 404 on a missing page; links
reports out/backlinks.

Claude-Session: https://claude.ai/code/session_01XMxbhmAkxxVV8CFWCZDBaG

* feat(web): delete a knowledge base from the settings sheet (type-name confirm)

Adds a Danger Zone section to KbSettingsSheet: a type-the-name confirmation
gates deleteKb() (POST /api/v1/kb/delete); on success KbDetail navigates back
to the KB list. New deleteKb client + kbSettings danger* keys + common cancel
(zh + en, identical key sets).

Claude-Session: https://claude.ai/code/session_01XMxbhmAkxxVV8CFWCZDBaG

* feat(web): in-reader page edit + delete with impact preview (F2/F3)

For an open concept/entity page, the reader header now shows Edit and Delete:
- Delete: a dry-run first (deletePage dryRun) surfaces the backlink pages whose
  [[links]] will demote to plain text; a red confirm card lists them, then the
  real delete closes the page and refreshes the inventory.
- Edit: a body textarea (frontmatter is preserved server-side), Save via
  PUT /api/v1/page, a links panel (getPageLinks: out/backlinks), a "recompile
  may overwrite" note, and a toast listing any dead links demoted to text.

Adds the deletePage/editPage/getPageLinks wiki clients and kb:pageOps.* locale
keys (zh + en, identical key sets). Build green (i18n guard + tsc + vite).

Claude-Session: https://claude.ai/code/session_01XMxbhmAkxxVV8CFWCZDBaG

* fix(workbench): address code-review findings (content-mgmt correctness + concurrency)

Backend (page_ops / kb_admin / lint):
- delete/edit now do their read-modify-write INSIDE the KB ingest lock and
  re-check the page exists under it: no stale backlink snapshot, no resurrecting
  a concurrently-deleted page, no 500 on a racing unlink (missing_ok=True). [#4,#5,#6]
- delete_kb serializes rmtree under the ingest lock (was fully unlocked). Moved
  delete_kb/unregister_kb into a new openkb/kb_admin.py so config.py drops back
  under the per-file line gate (751 lines). [#1,#15]
- page deletion no longer scans or rewrites lint's _EXCLUDED_FILES
  (AGENTS.md/SCHEMA.md/log.md); fix_broken_links(restrict_to) also refuses them,
  which protects `openkb remove` too. [#2]
- edit_wiki_page treats `content` as the body verbatim — no naive frontmatter
  split that silently dropped a body opening with a '---' block. [#3]
- delete uses the real on-disk stem for the exact-match index-entry removal
  (case-insensitive FS), and adds index.md to the demotion set so a [[target]]
  embedded in another entry's brief no longer dangles. [#7,#9]

API / CLI:
- delete-KB endpoint & CLI can now clean a ghost registry entry (registered name
  whose directory was removed by hand), catch delete_kb's ValueError -> 400, and
  the endpoint is typed `-> KbDeleteResponse`. [#8,#13,E5]

Frontend (KbDetail):
- Deleting a KB dispatches `openkb:reload-kbs` so the sidebar refreshes
  immediately (was stale until a manual reload) — live-test fix.
- Removed the dead `pageReloadSeq` path (editPage always returns content). [#14]

Tests: excluded-doc-untouched, '---'-body preservation, ghost-KB unregister;
non-KB-target test updated to the new resolution model. Backend 1234 passed,
mypy/ruff clean, frontend build green.

Not changed (intentional): pages_linking_to is NOT folded into build_graph (that
excludes explorations/ backlinks) [#10]; the preview+confirm delete inherently
scans per request [#11]; EDITABLE_SECTIONS stays duplicated in the frontend,
which can't import the Python constant [#12].

Claude-Session: https://claude.ai/code/session_01XMxbhmAkxxVV8CFWCZDBaG

* feat(workbench): allow editing summary pages (edit ⊇ delete section sets)

Summaries are compiled markdown like concepts/entities (a recompile regenerates
them), so they are now editable via PUT /api/v1/page. They remain NOT
independently deletable — a summary is removed by deleting its source document.

Split page_ops into EDITABLE_SECTIONS (concepts/entities/summaries) vs
DELETABLE_SECTIONS (concepts/entities); validate_page_ref takes the allowlist.
Frontend ReaderBody gates Edit on canEdit (incl. summaries), Delete on canDelete.

Test: summary editable (frontmatter preserved) + summary delete rejected (400).

Claude-Session: https://claude.ai/code/session_01XMxbhmAkxxVV8CFWCZDBaG

* fix(api): cross-platform delete_kb + endpoint OSError handling (xhigh review #1,#2)

- delete_kb no longer holds the ingest lock DURING rmtree (the lock file lives
  inside the KB dir; Windows cannot delete an open file — the prior review-fix
  regressed this). It now takes the lock as a BARRIER (drain + wait out any
  in-flight mutation), releases it, re-checks existence, then rmtrees. [#1]
- delete-KB endpoint maps FileNotFoundError to an idempotent success (a
  concurrent delete already removed the tree) and other OSError to a clean 500
  with a message, instead of an uncaught 500 stack trace. [#2]

Tests: endpoint OSError -> clean 500, FileNotFoundError -> 200 deleted.

Claude-Session: https://claude.ai/code/session_01XMxbhmAkxxVV8CFWCZDBaG
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

2 participants