- Add POST /api/v1/file/:file_uuid/cluster-agent endpoint for on-demand face clustering
- Fix OCR chunks being labeled as ASRX: use ChunkType.as_str() instead of {:?}
- Rule 1 now deletes old chunks before re-inserting to avoid stale data
- Add fallback face_traced.json when store_traced_faces.py fails
- ingestion_complete now handles status='error' to unblock jobs
- VLM describe tool now uses trace_id with pre-extracted face crops
- Update system prompt to prioritize smart_search over find_file
- Add source prefix ([OCR], [ASRX], [ASRX+OCR]) to exec_smart_search results
- Users can now see chunk content with source indicators
- Add [OCR], [ASRX], [ASRX+OCR] prefix to text_content
- Add content field to SemanticSearchResult struct
- Update SQL queries to include content field
- Helps users distinguish the source of search results
- Fix OCR confidence threshold: 0.5 → 0.2
- Fix fetch_ocr_texts frame calculation (use start_frame from DB)
- Fix text_match filter to use text_content instead of summary
- Process all merged results (not just top 30) to include keyword results
- Add logging for keyword search debugging
Fixes issue where OCR text like 'AUDIO MONITORING' and 'thunderbolt'
could not be found via keyword search.
- agent_search.rs: text_region -> text_trace, removed skin_tone_trace
(never built, misleading LLM agents)
- tkg.rs + scan.rs: hand_object -> HAND_OBJECT for consistency with
all other edge types (CO_OCCURS_WITH, SPEAKS_AS, etc.)
- rule2_ingest.rs: added LIP_SYNC and HAND_OBJECT to edge_type
priority list so these edges generate relationship chunks
- Collect edges in Vec first, then batch insert in chunks of 100
- Uses sqlx::query_builder::QueryBuilder for efficient batch INSERT
- Reduces SQL round-trips from O(edges) to O(edges/100)
- Maintains ON CONFLICT DO UPDATE semantics
- scroll_face_points now retries up to 3 times on failure
- Exponential backoff: 1s, 2s, 4s between attempts
- Logs retry attempts for debugging
- Prevents silent failures when Qdrant is temporarily unavailable
- Create src/core/db/schema.rs with t() and table_name() functions
- Remove duplicate t() from tkg.rs and trace_agent_api.rs
- postgres_db.rs uses schema::table_name() from shared module
- Add build_node_id_map() to fetch all node IDs once into HashMap
- Pass node_id_map to all 6 edge builders
- build_co_occurrence_edges: uses map lookup instead of per-face SELECT
- build_speaker_face_edges: uses map lookup instead of per-trace SELECT
- Remove redundant Qdrant scroll in build_co_occurrence_edges
- Expected: 90%+ reduction in SQL queries during edge building
- Created src/core/tkg/ module with service, log, and models
- Added tkg_operation_log table for tracking all TKG operations
- TkgService: build, rebuild (with force option), delete, get_operations
- Updated rebuild endpoint with force parameter
- Added GET /api/v1/file/:file_uuid/tkg for operation history
- Added DELETE /api/v1/file/:file_uuid/tkg for TKG deletion
- TKG rebuild now always triggers Rule 2 (even with 0 edges)
- Full audit trail for all TKG operations (create/update/delete/rebuild)
When ASRX falls back to ASR segments, end_frame may be 0.
The FPS calculation now handles this case correctly by checking
both end_frame > 0 and end_time > 0 before dividing.
This prevents division by zero and incorrect FPS values when
processing videos with ASRX fallback segments.
Previously Rule 1 only created chunks from ASRX segments, merging OCR
text where frame ranges overlapped. OCR text that didn't overlap with
any ASRX segment was ignored.
Now Rule 1 has two phases:
1. Process ASRX segments (merge OCR where overlapping) - existing behavior
2. Create chunks for OCR-only text (frames not covered by ASRX)
OCR-only chunks are grouped by consecutive frames (within 5 frames)
to avoid creating too many single-frame chunks.
Example: ASRX 819 + OCR-only 4 = 823 sentence chunks
- Skip chunks where both ASRX text and OCR text are empty
- Use count-based chunk_id instead of index to avoid gaps
- This ensures PostgreSQL and Qdrant chunk counts match
- Added text_content field to SearchResult and SemanticSearchResult
- Added get_chunk_by_id_no_embedding for keyword results without embedding requirement
- Fixed search_bm25 to use position-based ranking for CJK/Korean content
- Fixed sqlx column mapping with explicit alias
- Skip text_match filter for keyword-only results
- Use text_content as fallback when summary is empty
- Removed trace_chunks field from PostgresStats struct
- Removed trace_chunks query from get_file_stats and get_ingestion_status
- Fixed OCR fetch_ocr_texts to compute frames from start_time*FPS
- Updated scan.rs to use separate count_nodes/count_edges functions
- tkg_nodes has no edge_type column, query was failing silently
- Split into count_nodes(node_type) and count_edges(edge_type)
- Fixed text_region → text_trace node type name
- Also: OCR frame fix in rule1 (end_frame computed from end_time+FPS)
- get_file_identities: UNION face_detections + file_identities
- list_identities: add file_bindings from file_identities table
- Add back /api/v1/traces/unassigned route
- Total count query now includes file_identities
Frontend can now:
- Filter pending identities by file_uuid
- Filter pending faces (unassigned traces) by file_uuid
- skin_tone is a person attribute (like height), not trace attribute
- Remove build_skin_tone_trace_nodes function
- Remove skin_tone_trace_nodes from TkgResult and API response
- Remove skin_tone_trace from documentation tables
- Fix trace_id type mismatch (INT4 vs i64) with explicit ::bigint cast
- Change build_face_track_nodes to use from_pg version
- Add skin_tone_trace_nodes to API response
- Add #[derive(Serialize)] to TkgResult
- Fix Unicode panic in text label truncation
- Add push_existing_embeddings.py script
- Add Queued variant to VideoStatus enum
- Trigger sets videos.status='queued' instead of staying 'pending'
- Worker sets videos.status='processing' on pickup
- list_monitor_jobs_by_status ORDER BY created_at ASC (FIFO)
- queue_position counts both 'pending' and 'queued' jobs
- Identity agent: per-face max matching, multi-round with derived
seeds from high-confidence faces, angle diversity filter (cosine sim < 0.90)
- Pending person API: POST /file/:file_uuid/pending-person
+ GET /file/:file_uuid/pending-persons with status=pending, source=manual
- Update API docs (07_identity.md)
- Remove Rule 3 (Scene Chunking) from worker auto-trigger
- Remove rule3_ingest.rs and related imports
- Remove Story/Caption from playground module parsing
- Clean up scan.rs Rule 3 display
- Fix ASRX field name conversion (start_time -> start)
Reason: Story/5W1H/Scene accuracy too poor - will redesign later
- Problem: compact=p=0:nk=1 outputs pipe-delimited format without pts_time=
- Fix: default=nk=0 outputs pts_time=XXX format that parser can match
- Result: Charade scene detection from 1 scene -> 833 scenes (correct)
- pre_chunks: add chunk_type, text_content columns; drop NOT NULL on
coordinate_type/coordinate_index (INSERT statements reference these
columns but CREATE TABLE was missing them)
- run_migrations: add ALTER TABLE for existing databases
- extract_movie_name: filter noise words (youtube, fps, 24fps, 1080p,
pure digits) so 'Charade_YouTube_24fps' → 'Charade'
- run-server-3002.sh: add companion worker startup (matching 3003 script)
- ProcessorType::all(): remove MediaPipe, Appearance, Story (mediapipe replaced by Swift)
- files.rs auto-pipeline: fix order to cut,asr,asrx,yolo,ocr,face,pose (was missing asr)
- postgres_db.rs run_migrations(): rewrite to auto-create all 38 tables idempotently