fix: face group name read consistency, sync_file_status fix, cleanup ghost records, identity_agent replaced with face_dedup
- get_face_groups_handler: COALESCE(tp.name, tn.label) for name consistency - sync_file_status: compare JSON vs pre_chunks (not chunk table) - face consistency: compare frames.len() not total_faces - cleanup 2 ghost records with NULL file_name/file_path - replace identity_agent with face_dedup in pipeline stages - remove identity_agent_api.rs and all references - update required_processors to match actual processors - update AGENTS.md with team responsibilities - add Studio pipeline changes documentation
This commit is contained in:
@@ -0,0 +1,148 @@
|
||||
---
|
||||
title: Always-Produce Processing Contract
|
||||
version: 1.0
|
||||
date: 2026-07-24
|
||||
author: OpenCode
|
||||
status: approved
|
||||
---
|
||||
|
||||
# Always-Produce Processing Contract
|
||||
|
||||
## Scope
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| Scope | All frame-based processors (face, pose, appearance, face_cluster, face_trace, etc.) |
|
||||
| Status | Approved |
|
||||
| Applies to | Python processors + Rust Worker |
|
||||
| Related docs | `DESIGN/Processor_Module_V1.0.md`, `DESIGN/Redis_Progress_Reporting_V1.0.md`, `DESIGN/Worker_Health_Check_Mechanism.md` |
|
||||
|
||||
## 1. Frame-Scan Model
|
||||
|
||||
Video processing is fundamentally frame-based: a processor scans from frame 0 to the last frame.
|
||||
|
||||
```
|
||||
Scan start → frame 0 → frame 1 → ... → frame N → scan complete
|
||||
↓ ↓ ↓ ↓
|
||||
Redis Redis Redis {uuid}.{p}.json
|
||||
progress progress progress (final record)
|
||||
```
|
||||
|
||||
### Key Rules
|
||||
|
||||
1. **Progress** = which frame has been scanned so far (`current_frame / total_frames`)
|
||||
2. **Complete** = scanned to the last frame (proved by `.json` existing)
|
||||
3. **Result** = always written, even if 0 detections found
|
||||
|
||||
## 2. Always-Produce Rule
|
||||
|
||||
### Principle
|
||||
|
||||
> Every processor MUST write its `{uuid}.{processor}.json` output file after completing its scan, **regardless of whether any results were found**.
|
||||
|
||||
### Rationale
|
||||
|
||||
The `.json` file serves dual purpose:
|
||||
- **Proof of completion**: Worker uses `output_path.exists()` (line 580 of `job_worker.rs`) to skip already-finished processors
|
||||
- **Downstream dependency**: Subsequent processors check this file for input
|
||||
|
||||
Without the Always-Produce rule:
|
||||
- Zero-result processors leave no `.json` → Worker retries infinitely → deadlock
|
||||
- Stuck jobs block downstream stages (Rule 1/2/3 ingestion, TKG build)
|
||||
|
||||
### Format
|
||||
|
||||
All processor JSON outputs MUST include:
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "has_faces" | "no_faces" | "no_face_json" | "no_embeddings" | "success" | "error_*",
|
||||
"file_uuid": "<uuid>",
|
||||
...processor-specific fields (empty arrays when zero results)
|
||||
}
|
||||
```
|
||||
|
||||
Example — face cluster with no faces:
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "no_faces",
|
||||
"file_uuid": "9781de6d...",
|
||||
"clusters": [],
|
||||
"frames": []
|
||||
}
|
||||
```
|
||||
|
||||
### Processor Checklist
|
||||
|
||||
| Processor | Always-Produce? | Status field on 0 result |
|
||||
|-----------|----------------|--------------------------|
|
||||
| `face.py` | ✅ Yes | `"no_faces"` |
|
||||
| `store_traced_faces.py` | ✅ Yes (already writes) | `"no_faces"` |
|
||||
| `fast_face_clustering_processor.py` | ❌ **FIX NEEDED** | Early returns, no file written |
|
||||
| `pose_processor*.py` | ✅ Yes | `"no_faces"` |
|
||||
| `appearance_processor*.py` | ✅ Yes | `"no_faces"` |
|
||||
|
||||
## 3. Redis Progress During Scan
|
||||
|
||||
### Purpose
|
||||
|
||||
Live frame progress is published to Redis so the QC modal can display real-time status ("scanning frame 1234/5678").
|
||||
|
||||
### Mechanism
|
||||
|
||||
Use `redis_publisher.py` (`RedisPublisher` class) which publishes to Redis channel `{prefix}progress:{uuid}`:
|
||||
|
||||
```python
|
||||
from redis_publisher import RedisPublisher
|
||||
|
||||
pub = RedisPublisher(file_uuid)
|
||||
|
||||
# During scan, per batch:
|
||||
pub.progress("face_cluster", current_frame, total_frames, f"Scanning frame {current_frame}")
|
||||
|
||||
# On completion:
|
||||
pub.complete("face_cluster", f"Done: {cluster_count} clusters")
|
||||
```
|
||||
|
||||
### Frequency
|
||||
|
||||
- **Frame-based processors**: publish every N frames (batch/buffer flush)
|
||||
- **Non-frame processors** (e.g., clustering): publish at meaningful milestones
|
||||
|
||||
## 4. Worker Heartbeat
|
||||
|
||||
### Problem
|
||||
|
||||
`health.rs` currently uses `check_process_running("worker")` which relies on `ps aux | grep momentry.*worker`. This is unreliable:
|
||||
- Zombie processes show as "running"
|
||||
- Stale matches from unrelated processes
|
||||
|
||||
### Fix
|
||||
|
||||
Worker writes a Redis HMSET `{prefix}health` with EXPIRE = `3 × poll_interval_secs` (default: 15s) in every `poll_and_process()` cycle.
|
||||
|
||||
Health endpoint checks:
|
||||
1. Redis key `{prefix}health` exists
|
||||
2. Key has remaining TTL > 0
|
||||
3. Key's `status` field is `"healthy"` or `"throttled"`
|
||||
|
||||
If Redis key missing or expired → `worker_alive: false`.
|
||||
|
||||
## 5. Implementation Plan
|
||||
|
||||
| Step | File | Change |
|
||||
|------|------|--------|
|
||||
| 1 | `fast_face_clustering_processor.py` | Always-Produce for 3 early returns + Redis progress |
|
||||
| 2 | `store_traced_faces.py` | Add Redis progress (optional) |
|
||||
| 3 | `job_worker.rs` | Add EXPIRE after health HMSET |
|
||||
| 4 | `health.rs` | Replace `check_process_running("worker")` with Redis TTL check |
|
||||
| 5 | `processing.rs` | (Optional) Reject trigger if Worker not alive |
|
||||
|
||||
---
|
||||
|
||||
## Version History
|
||||
|
||||
| Version | Date | Author | Changes |
|
||||
|---------|------|--------|---------|
|
||||
| 1.0 | 2026-07-24 | OpenCode | Initial specification |
|
||||
@@ -0,0 +1,341 @@
|
||||
---
|
||||
title: Face Tracking Pipeline Structure
|
||||
version: 1.0
|
||||
date: 2026-07-22
|
||||
author: OpenCode
|
||||
status: Active
|
||||
---
|
||||
|
||||
# Face Tracking Pipeline — Structure Design
|
||||
|
||||
## Overview
|
||||
|
||||
```
|
||||
Video
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────┐
|
||||
│ Stage 1: Face Detection │ face_processor.py
|
||||
│ swift_face (Apple Vision) │ → {uuid}.face.json
|
||||
│ CoreML FaceNet embedding │ → Qdrant _faces (initial)
|
||||
└──────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────┐
|
||||
│ Stage 2: Face Tracking │ store_traced_faces.py
|
||||
│ face_tracker.py (IoU) │ → {uuid}.face_traced.json
|
||||
│ trace_id assignment │ → Qdrant _faces (trace_id update)
|
||||
└──────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────┐
|
||||
│ Stage 3: Trace Profile │ backfill_trace_profiles.py
|
||||
│ Qdrant _faces 分組 │ → output/{uuid}/trace_{N}/
|
||||
│ key_frame + key_face │ trace_profile.json
|
||||
└──────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────┐
|
||||
│ Stage 4: TKG Nodes │ tkg.rs
|
||||
│ Qdrant _faces → trace_id │ → tkg_nodes (face_track, etc.)
|
||||
└──────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Stage 1: Face Detection
|
||||
|
||||
**Script**: `scripts/face_processor.py`
|
||||
|
||||
### Flow
|
||||
|
||||
1. `swift_face` (Swift/Apple Vision/ANE) → bbox detection per sampled frame
|
||||
2. `cv2` opens video, crops face from bbox
|
||||
3. CoreML FaceNet → 512D embedding per face
|
||||
4. Output: `{uuid}.face.json`
|
||||
5. Push embeddings to Qdrant `_faces` collection
|
||||
|
||||
### Output Format: `{uuid}.face.json`
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "has_faces",
|
||||
"frame_count": 563,
|
||||
"fps": 29.97,
|
||||
"total_faces": 1200,
|
||||
"frames": [
|
||||
{
|
||||
"frame": 743,
|
||||
"timestamp": 24.78,
|
||||
"faces": [
|
||||
{
|
||||
"x": 892, // int, pixel
|
||||
"y": 313, // int, pixel
|
||||
"width": 78, // int, pixel
|
||||
"height": 78, // int, pixel
|
||||
"confidence": 0.733,
|
||||
"pose_angle": { "angle": "frontal", "roll": 0.77, "yaw": -1.24, "pitch": 0.23 },
|
||||
"landmarks": { "right_eye": [...], "nose": [...], "left_eye": [...] },
|
||||
"lips": { "inner_lips": [...], "outer_lips": [...] }
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Key points**:
|
||||
- bbox is **pixel integer** from Apple Vision, never modified
|
||||
- face.json uses **list format** (not dict)
|
||||
- Sampling at ~8Hz (`sample_interval = round(fps / 8)`)
|
||||
|
||||
### Qdrant Initial Push
|
||||
|
||||
`push_face_embeddings_batch()` in `qdrant_faces.py`:
|
||||
|
||||
```python
|
||||
payload = {
|
||||
"file_uuid": file_uuid,
|
||||
"frame": frame_num,
|
||||
"trace_id": face_idx, # ⚠️ frame-internal index (0, 1, 2...), NOT tracking trace_id
|
||||
"bbox": {"x": x, "y": y, "width": w, "height": h}, # int pixel
|
||||
"confidence": 0.5,
|
||||
"identity_id": None,
|
||||
"identity_uuid": None,
|
||||
"stranger_id": None,
|
||||
}
|
||||
```
|
||||
|
||||
**Important**: `trace_id` at this stage is `face_idx` (index within the frame), used only as a temporary placeholder. It gets overwritten in Stage 2.
|
||||
|
||||
---
|
||||
|
||||
## Stage 2: Face Tracking
|
||||
|
||||
**Scripts**: `scripts/store_traced_faces.py` → `scripts/utils/face_tracker.py`
|
||||
|
||||
### Trigger
|
||||
|
||||
`job_worker.rs` P2 trigger (line ~1877): after face + asrx processors complete.
|
||||
|
||||
```rust
|
||||
tokio::spawn(async move {
|
||||
executor.run("store_traced_faces.py", &["--file-uuid", &uuid], ...)
|
||||
});
|
||||
```
|
||||
|
||||
Skip if `{uuid}.face_traced.json` already exists.
|
||||
|
||||
### Flow
|
||||
|
||||
1. `store_traced_faces.py` reads `{uuid}.face.json`
|
||||
2. Converts face.json from list to dict format (frame_num_str → {frame_number, time_seconds, faces})
|
||||
3. Loads cut boundaries from `{uuid}.cut.json` (if exists)
|
||||
4. Calls `face_tracker.track_faces(face_data, use_embedding=False, cut_boundaries=...)`
|
||||
5. Writes `{uuid}.face_traced.json`
|
||||
6. Calls `update_trace_ids(file_uuid, trace_mapping)` to update Qdrant
|
||||
|
||||
### `face_tracker.py:track_faces()`
|
||||
|
||||
**Algorithm** (IoU-only, no embedding):
|
||||
|
||||
```
|
||||
For each frame (sorted):
|
||||
For each face in current frame:
|
||||
Match against previous frame faces:
|
||||
- Calculate IoU
|
||||
- Calculate bbox center distance
|
||||
- Reject if area ratio > 5x (different zoom level)
|
||||
- Reject if at-edge → not-at-edge transition (person exited)
|
||||
If match found → same trace_id as matched face
|
||||
If no match → new trace_id (next_trace_id++)
|
||||
Scene cut boundary between frames → force all new traces
|
||||
```
|
||||
|
||||
**Matching conditions** (IoU-only mode):
|
||||
- IoU > 0.5 AND IoU > 0.35 + distance < 100px → match
|
||||
- IoU > 0.5 + similarity > 0.65 → match (similarity not used but condition exists)
|
||||
- similarity > 0.85 → match (not used in IoU-only mode)
|
||||
- Scene cut boundary → all new traces
|
||||
|
||||
### Output Format: `{uuid}.face_traced.json`
|
||||
|
||||
Same structure as face.json, but:
|
||||
- Format converted to **dict** (`frames[str(frame_num)]` → face data)
|
||||
- Each face gains `trace_id` field (integer)
|
||||
- Top-level `traces` dict with per-trace statistics
|
||||
- `metadata.tracking_method = "iou_only"`
|
||||
- `metadata.traced_at = ISO timestamp`
|
||||
|
||||
```json
|
||||
{
|
||||
"metadata": {
|
||||
"fps": 29.97,
|
||||
"total_frames": 43977,
|
||||
"tracking_method": "iou_only",
|
||||
"trace_stats": {
|
||||
"total_traces": 107,
|
||||
"active_traces": 107,
|
||||
"long_traces": 95
|
||||
}
|
||||
},
|
||||
"frames": {
|
||||
"743": {
|
||||
"frame_number": 743,
|
||||
"faces": [
|
||||
{ "x": 892, "y": 313, "width": 78, "height": 78, "trace_id": 0, ... }
|
||||
]
|
||||
}
|
||||
},
|
||||
"traces": {
|
||||
"0": {
|
||||
"trace_id": 0,
|
||||
"start_frame": 743,
|
||||
"end_frame": 783,
|
||||
"duration_frames": 41,
|
||||
"total_appearances": 11,
|
||||
"avg_confidence": 0.72
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Qdrant Trace Update
|
||||
|
||||
`update_trace_ids()` in `qdrant_faces.py`:
|
||||
|
||||
1. Scroll all Qdrant `_faces` points for `file_uuid` (with vector + payload)
|
||||
2. For each point, build `bbox_key = f"{bbox.x}_{bbox.y}_{bbox.width}_{bbox.height}"`
|
||||
3. Look up `trace_mapping[frame][bbox_key]` from face_traced.json
|
||||
4. If match found → set `payload["trace_id"] = real_trace_id`
|
||||
5. PUT updated points back to Qdrant
|
||||
|
||||
**Matching key**: `frame` + `bbox_key` (pixel integer string)
|
||||
|
||||
---
|
||||
|
||||
## Stage 3: Trace Profile
|
||||
|
||||
**Script**: `scripts/backfill_trace_profiles.py`
|
||||
|
||||
### Data Source
|
||||
|
||||
Qdrant `_faces` collection (source of truth for trace_id assignments).
|
||||
|
||||
### Flow
|
||||
|
||||
1. Scroll all `_faces` points for each `file_uuid` with `trace_id >= 0`
|
||||
2. Group by `(file_uuid, trace_id)`
|
||||
3. For each group:
|
||||
- `frame_count` = count of points
|
||||
- `start_frame` = min(frame)
|
||||
- `end_frame` = max(frame)
|
||||
- `representative_frame` = frame with max(confidence)
|
||||
- `representative_bbox` = bbox at representative frame
|
||||
4. Extract `key_frame.jpg` via ffmpeg at representative frame
|
||||
5. Crop `key_face.jpg` from key_frame using representative bbox
|
||||
6. Write `output/{uuid}/trace_{N}/trace_profile.json`
|
||||
|
||||
### Output: `output/{uuid}/trace_{N}/trace_profile.json`
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "1.0",
|
||||
"file_uuid": "d8acb03870f0cc9b14e01f14a7bf24d6",
|
||||
"trace_id": 37,
|
||||
"label": "",
|
||||
"frame_count": 38,
|
||||
"start_frame": 1859,
|
||||
"end_frame": 2100,
|
||||
"avg_confidence": 0.754,
|
||||
"key_frame": "key_frame.jpg",
|
||||
"key_face": "key_face.jpg",
|
||||
"status": "pending"
|
||||
}
|
||||
```
|
||||
|
||||
### File Layout
|
||||
|
||||
```
|
||||
output/{uuid}/
|
||||
trace_0/
|
||||
trace_profile.json
|
||||
key_frame.jpg
|
||||
key_face.jpg
|
||||
trace_1/
|
||||
trace_profile.json
|
||||
key_frame.jpg
|
||||
key_face.jpg
|
||||
...
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Stage 4: TKG Node Construction
|
||||
|
||||
**File**: `src/core/processor/tkg.rs`
|
||||
|
||||
Reads trace_id from Qdrant `_faces` payload to build knowledge graph nodes:
|
||||
- `face_track` nodes: one per trace
|
||||
- `gaze_track`, `lip_track`: linked to face_track via frame alignment
|
||||
- `co_occurrence` edges: traces that appear in same frame
|
||||
|
||||
---
|
||||
|
||||
## Qdrant `_faces` Collection Schema
|
||||
|
||||
| Field | Type | Description |
|
||||
|-------|------|-------------|
|
||||
| `file_uuid` | string | Video file identifier |
|
||||
| `frame` | int | Video frame number (absolute, not sampled) |
|
||||
| `trace_id` | int | Face tracking ID (set by Stage 2) |
|
||||
| `bbox` | `{x, y, width, height}` | Pixel integer coordinates |
|
||||
| `confidence` | float | Detection confidence |
|
||||
| `identity_id` | int? | Identity binding (set by identity agent) |
|
||||
| `identity_uuid` | string? | Identity UUID |
|
||||
| `stranger_id` | int? | Stranger classification |
|
||||
|
||||
**Point ID**: `generate_point_id(file_uuid, frame, face_idx)` — deterministic hash.
|
||||
|
||||
---
|
||||
|
||||
## Known Issues
|
||||
|
||||
### bfba056f5021e2404b0870cc0b1fa851
|
||||
|
||||
- **Qdrant**: trace_id = 0,1,2 (face_idx, never updated)
|
||||
- **face_traced.json**: trace_id = 0-8209 (8210 traces, iou_only)
|
||||
- **Root cause**: `face_processor.py` re-ran after `store_traced_faces.py`, pushing fresh embeddings with `trace_id=face_idx`, overwriting the updated trace_ids
|
||||
- **Other 12 files**: all correct
|
||||
|
||||
### `update_trace_ids` bbox matching
|
||||
|
||||
Matching is by exact `frame` + `bbox_key` string (`x_y_width_height`). Since bbox is pixel integer from the same source, values are identical across face_traced.json and Qdrant. Mismatch only occurs when face_processor.py re-runs and generates different detection results.
|
||||
|
||||
---
|
||||
|
||||
## File Inventory (2026-07-22)
|
||||
|
||||
| file_uuid | traces (Qdrant) | traces (face_traced) | status |
|
||||
|-----------|-----------------|----------------------|--------|
|
||||
| 30affad3... | 52 | 53 | ✅ |
|
||||
| 31a6b821... | 31 | 36 | ⚠️ minor mismatch |
|
||||
| 352cf73a... | 16 | 25 | ⚠️ minor mismatch |
|
||||
| 57bd7e43... | 3 | 4 | ✅ |
|
||||
| 5e5f3de8... | 21 | 22 | ✅ |
|
||||
| 84d838f2... | 88 | 89 | ✅ |
|
||||
| 88e72467... | 18 | 19 | ✅ |
|
||||
| 9cbeb112... | 9 | 17 | ⚠️ minor mismatch |
|
||||
| bfba056f... | 15 | 8210 | ❌ face_idx not updated |
|
||||
| c0a9dc37... | 77 | 78 | ✅ |
|
||||
| c36f3568... | 5601 | 5616 | ⚠️ minor mismatch |
|
||||
| d8acb038... | 106 | 107 | ✅ |
|
||||
| fbd82072... | 12 | 13 | ✅ |
|
||||
|
||||
---
|
||||
|
||||
## Version History
|
||||
|
||||
| Version | Date | Changes |
|
||||
|---------|------|---------|
|
||||
| 1.0 | 2026-07-22 | Initial document: face detection → tracking → Qdrant → TKG pipeline structure |
|
||||
@@ -1,198 +1,513 @@
|
||||
---
|
||||
document_type: "design_doc"
|
||||
service: "MOMENTRY_CORE"
|
||||
title: "File Lifecycle — Pre-Processing & Registration"
|
||||
version: "V1.2"
|
||||
date: "2026-05-15"
|
||||
author: "M5"
|
||||
status: "draft"
|
||||
title: File Lifecycle Architecture
|
||||
version: 1.0
|
||||
date: 2026-07-22
|
||||
author: OpenCode
|
||||
status: Active
|
||||
scope: File processing pipeline — stages, verification, rebuild
|
||||
---
|
||||
|
||||
# File Lifecycle — Pre-Processing & Registration
|
||||
# File Lifecycle Architecture V1.0
|
||||
|
||||
| Item | Value |
|
||||
|------|-------|
|
||||
| Scope | All managed file types (video, image, document, spreadsheet, presentation) |
|
||||
| Status | Draft |
|
||||
| Applies to | Pre-process API (explicit) + Register API |
|
||||
| Key concept | Two-phase flow: birth certificate (`.pre.json`) → civil registration (DB INSERT) |
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| Scope | Complete file processing lifecycle |
|
||||
| Status | Active |
|
||||
| Applies to | Pipeline stages, progress tracking, verification, rebuild |
|
||||
| Related | `FILE_PROFILE_V1.0.md`, `FACE_TRACKING_PIPELINE_V1.0.md` |
|
||||
|
||||
> **Applicable to all managed file types**: video, image, document (pdf, docx, pages, key, numbers), spreadsheet, presentation, and any other file registered in the system. The pre-processor registers any file type found by the watcher. ffprobe is used when applicable; files that ffprobe cannot parse receive minimal filesystem metadata as a fallback.
|
||||
---
|
||||
|
||||
## Metaphor
|
||||
## 1. Overview
|
||||
|
||||
Every registered video file passes through a deterministic pipeline of stages.
|
||||
Each stage must produce a `.json` (or `.jpg`) artifact on disk.
|
||||
This enables:
|
||||
- **Verification**: Check pipeline completeness by inspecting artifact existence
|
||||
- **Rebuild**: Re-run any stage from its input artifacts without re-running the entire pipeline
|
||||
- **Progress tracking**: Two-layer display (high-level summary + expandable sub-stages)
|
||||
|
||||
### Design Principles
|
||||
|
||||
1. **Every stage has a `.json` output** — no silent DB-only writes
|
||||
2. **Any stage can be rebuilt** from its input artifacts
|
||||
3. **Frontend reads stages from API** — not hardcoded
|
||||
4. **Verification is disk-first** — check `.json` exists, then validate content, then check DB/Qdrant consistency
|
||||
5. **Processors are not modified** — this document defines tracking/verification/rebuild only
|
||||
|
||||
---
|
||||
|
||||
## 2. Stage Architecture
|
||||
|
||||
### 2.1 High-Level Stages (6)
|
||||
|
||||
| # | Stage | Weight | Sub-Stages | Description |
|
||||
|---|-------|--------|------------|-------------|
|
||||
| S0 | Register | 5% | 4 | File metadata + audio track + key frame extraction |
|
||||
| S1 | Processors | 40% | 8 | Individual processor execution |
|
||||
| S2 | Post-Process | 20% | 4 | Face trace, Rule1, Vectorize, Identity Agent |
|
||||
| S3 | TKG Build | 20% | 2 | Temporal Knowledge Graph nodes + edges |
|
||||
| S4 | Rule2 | 10% | 1 | Relationship chunk ingestion |
|
||||
| S5 | Complete | 5% | 1 | Final status update |
|
||||
|
||||
### 2.2 Sub-Stages (15)
|
||||
|
||||
```
|
||||
SHA256 = DNA or fingerprint (immutable biometric identity)
|
||||
file mtime = birth moment (preserved by rsync across systems)
|
||||
birthday (file_uuid anchor) = mtime timestamp
|
||||
.pre.json = birth certificate
|
||||
POST /api/v1/files/register = civil registration
|
||||
status = registered = citizenship completed
|
||||
S0: Register (5%)
|
||||
├─ 0a: probe → probe.json
|
||||
├─ 0b: audio_track → DB: audio_track column (no disk artifact)
|
||||
├─ 0c: profile → profile.json
|
||||
└─ 0d: key_frame → key_frame.jpg
|
||||
|
||||
S1: Processors (40%)
|
||||
├─ 1a: cut → cut.json
|
||||
├─ 1b: asr → asr.json
|
||||
├─ 1c: asrx → asrx.json (depends: 1a + 1b)
|
||||
├─ 1d: ocr → ocr.json
|
||||
├─ 1e: face → face.json (+ Qdrant _faces initial)
|
||||
├─ 1f: pose → pose.json (depends: 1e)
|
||||
├─ 1g: appearance → appearance.json (depends: 1f)
|
||||
└─ 1h: face_dedup → face_cluster.json (depends: 1e) [OPTIONAL + MANUAL]
|
||||
|
||||
S2: Post-Process (20%)
|
||||
├─ 2a: face_trace → face_traced.json (+ Qdrant trace_id update)
|
||||
├─ 2b: rule1 → rule1.json (ASRX → sentence chunks)
|
||||
├─ 2c: vectorize → vectorize.json (embeddings → PG + Qdrant)
|
||||
└─ 2d: identity_agent → identity_agent.json (optional)
|
||||
|
||||
S3: TKG Build (20%)
|
||||
├─ 3a: tkg_nodes → tkg_nodes.json
|
||||
└─ 3b: tkg_edges → tkg_edges.json
|
||||
|
||||
S4: Rule2 (10%)
|
||||
└─ 4a: rule2 → rule2.json (relationship chunks)
|
||||
|
||||
S5: Complete (5%)
|
||||
└─ 5a: complete → status = "completed"
|
||||
```
|
||||
|
||||
## Two-Phase Flow
|
||||
|
||||
A file enters the system in two distinct phases:
|
||||
|
||||
| Phase | Action | Analogy | Automatic? | Status |
|
||||
|-------|--------|---------|:----------:|:------:|
|
||||
| **Birth** | Pre-process: SHA256 + probe + file_uuid | 出生 + 醫院開出生證明 | ✅ Watcher | `unregistered` |
|
||||
| **Citizenship** | Register: INSERT into DB | 戶政事務所登記 | ❌ User API | `registered` |
|
||||
|
||||
## Phase 1: Pre-Processing (Birth)
|
||||
|
||||
### Trigger
|
||||
|
||||
Pre-processing is triggered explicitly via the register API or a dedicated pre-process endpoint. It is NOT automatic — the watcher only detects new files without modifying them.
|
||||
|
||||
### Computation Steps
|
||||
### 2.3 Dependency Graph
|
||||
|
||||
```
|
||||
1. fs::metadata(path).modified()
|
||||
→ birthday = file modification time (mtime, RFC 3339; preserved by rsync -a across systems)
|
||||
|
||||
2. SHA256(full file, streaming 64KB chunks)
|
||||
→ content_hash = 512-bit hex string (file DNA / fingerprint)
|
||||
|
||||
3. ffprobe (or minimal fs metadata fallback for non-video)
|
||||
→ probe_json
|
||||
|
||||
4. compute_birth_uuid(mac, birthday, canonical_path, filename)
|
||||
→ file_uuid = SHA256(mac | birthday | path | filename)[0:32]
|
||||
|
||||
5. Write {OUTPUT_DIR}/{file_uuid}.pre.json
|
||||
S0 (Register)
|
||||
└─→ S1 (Processors)
|
||||
├─ 1a (CUT) ─────┐
|
||||
├─ 1b (ASR) ─────┤
|
||||
│ └─→ 1c (ASRX) ──→ 2b (Rule1)
|
||||
├─ 1d (OCR) ──────────────────────→ 3a (TKG Nodes)
|
||||
├─ 1e (Face) ──┬─→ 1f (Pose) ──→ 1g (Appearance) ──→ 3a
|
||||
│ ├─→ 1h (FaceDedup) [manual]
|
||||
│ └─→ 2a (Face Trace) ──→ 3a
|
||||
└─────────────────────────────────────→ 3a
|
||||
│
|
||||
S2: 2c (Vectorize) ←── DB chunks │
|
||||
S2: 2d (IdentityAgent) ←── face_clusters │
|
||||
↓
|
||||
3b (TKG Edges)
|
||||
│
|
||||
↓
|
||||
4a (Rule2)
|
||||
│
|
||||
↓
|
||||
5a (Complete)
|
||||
```
|
||||
|
||||
### Output: `.pre.json` Schema
|
||||
---
|
||||
|
||||
Stored alongside other processor outputs:
|
||||
## 3. I/O Specification
|
||||
|
||||
### 3.1 Register (S0)
|
||||
|
||||
| Sub-Stage | Input | Output Artifact | DB Tables | Qdrant |
|
||||
|-----------|-------|----------------|-----------|--------|
|
||||
| 0a: probe | video file on disk | `{uuid}.probe.json` | — | — |
|
||||
| 0b: audio_track | probe.json, video file | DB column only | videos.audio_track | — |
|
||||
| 0c: profile | probe.json | `{uuid}.profile.json` | videos (INSERT/UPDATE) | — |
|
||||
| 0d: key_frame | probe.json | `{uuid}.key_frame.jpg` | — | — |
|
||||
|
||||
**Audio Track Classification** (S0b):
|
||||
|
||||
| Classification | Condition | ASR Behavior |
|
||||
|----------------|-----------|--------------|
|
||||
| `no_audio` | No audio track in video | Skip ASR, output `{"status": "no_audio"}` |
|
||||
| `silent_audio` | Audio track exists but no speech detected | Skip ASR, output `{"status": "silent_audio"}` |
|
||||
| `music_only` | Audio with no speech (music/sound effects) | Skip ASR, output `{"status": "music_only"}` |
|
||||
| `speech_only` | Audio with speech only (≥30% speech ratio) | Run ASR normally |
|
||||
| `speech_with_music` | Speech with background music (<30% speech ratio) | Run ASR normally |
|
||||
|
||||
### 3.2 Processors (S1)
|
||||
|
||||
| Sub-Stage | Input Artifacts | Output Artifact | DB Tables | Qdrant |
|
||||
|-----------|----------------|----------------|-----------|--------|
|
||||
| 1a: cut | probe.json | `{uuid}.cut.json` + `{uuid}_scene_{n}.jpg` | processor_results | — |
|
||||
| 1b: asr | video file | `{uuid}.asr.json` | processor_results | — |
|
||||
| 1c: asrx | cut.json, asr.json | `{uuid}.asrx.json` | speaker_detections | — |
|
||||
| 1d: ocr | video file | `{uuid}.ocr.json` | processor_results | — |
|
||||
| 1e: face | video file | `{uuid}.face.json` | processor_results | `_faces` (initial push) |
|
||||
| 1f: pose | face.json, video file | `{uuid}.pose.json` | processor_results | — |
|
||||
| 1g: appearance | pose.json, video file | `{uuid}.appearance.json` | processor_results | — |
|
||||
| 1h: face_dedup | face.json | `{uuid}.face_cluster.json` | face_clusters | — |
|
||||
|
||||
**Note**: 1h (Face Deduplication) is currently `optional + manual`. It will be integrated into the automated pipeline after testing is complete.
|
||||
|
||||
**Scene Key Frames** (1a post-process):
|
||||
|
||||
After CUT completes, extracts the middle frame from each scene as `{uuid}_scene_{n}.jpg` for VLM analysis:
|
||||
|
||||
| Output | Purpose |
|
||||
|---------|---------|
|
||||
| `{uuid}_scene_1.jpg` | Representative frame from scene 1 |
|
||||
| `{uuid}_scene_2.jpg` | Representative frame from scene 2 |
|
||||
| ... | ... |
|
||||
|
||||
These key frames enable:
|
||||
- VLM scene understanding (caption, objects, actions)
|
||||
- Scene-level search and filtering
|
||||
- Thumbnail generation for scene navigation
|
||||
|
||||
### 3.3 Post-Process (S2)
|
||||
|
||||
| Sub-Stage | Input Artifacts | Output Artifact | DB Tables | Qdrant |
|
||||
|-----------|----------------|----------------|-----------|--------|
|
||||
| 2a: face_trace | face.json | `{uuid}.face_traced.json` | — | `_faces` (trace_id update) |
|
||||
| 2b: rule1 | asrx.json | `{uuid}.rule1.json` | chunk, pre_chunks | — |
|
||||
| 2c: vectorize | chunk (DB) | `{uuid}.vectorize.json` | chunk_vectors | main collection |
|
||||
| 2d: identity_agent | face_cluster.json | `{uuid}.identity_agent.json` | file_identities | — |
|
||||
|
||||
### 3.4 TKG Build (S3)
|
||||
|
||||
| Sub-Stage | Input Artifacts | Output Artifact | DB Tables | Qdrant |
|
||||
|-----------|----------------|----------------|-----------|--------|
|
||||
| 3a: tkg_nodes | All processor JSONs, trace profiles | `{uuid}.tkg_nodes.json` | tkg_nodes | — |
|
||||
| 3b: tkg_edges | tkg_nodes.json, asrx.json | `{uuid}.tkg_edges.json` | tkg_edges | — |
|
||||
|
||||
### 3.5 Rule2 (S4)
|
||||
|
||||
| Sub-Stage | Input Artifacts | Output Artifact | DB Tables | Qdrant |
|
||||
|-----------|----------------|----------------|-----------|--------|
|
||||
| 4a: rule2 | tkg_edges.json, chunk (DB) | `{uuid}.rule2.json` | chunk (relationship type) | main collection |
|
||||
|
||||
### 3.6 Complete (S5)
|
||||
|
||||
| Sub-Stage | Input | Output | DB Tables |
|
||||
|-----------|-------|--------|-----------|
|
||||
| 5a: complete | All above stages verified | status = "completed" | videos.status |
|
||||
|
||||
---
|
||||
|
||||
## 4. Verification
|
||||
|
||||
### 4.1 Verification Levels
|
||||
|
||||
Each sub-stage has three verification levels:
|
||||
|
||||
| Level | Check | Description |
|
||||
|-------|-------|-------------|
|
||||
| L1: Artifact exists | `{uuid}.{stage}.json` on disk | Required for all stages |
|
||||
| L2: Content valid | JSON parseable + non-empty array/object | Ensures output is usable |
|
||||
| L3: DB/Qdrant consistent | Row count > 0 or point count > 0 | Ensures data was written |
|
||||
|
||||
### 4.2 Verification Matrix
|
||||
|
||||
| Sub-Stage | L1 (exists) | L2 (valid) | L3 (DB/Qdrant) |
|
||||
|-----------|:-----------:|:----------:|:--------------:|
|
||||
| 0a: probe | `.probe.json` | non-empty | — |
|
||||
| 0b: profile | `.profile.json` | has file_uuid | videos row exists |
|
||||
| 0c: key_frame | `.key_frame.jpg` | file size > 0 | — |
|
||||
| 1a: cut | `.cut.json` | non-empty | processor_results > 0 |
|
||||
| 1b: asr | `.asr.json` | non-empty | processor_results > 0 |
|
||||
| 1c: asrx | `.asrx.json` | non-empty | speaker_detections > 0 |
|
||||
| 1d: ocr | `.ocr.json` | non-empty | processor_results > 0 |
|
||||
| 1e: face | `.face.json` | non-empty | Qdrant `_faces` > 0 |
|
||||
| 1f: pose | `.pose.json` | non-empty | processor_results > 0 |
|
||||
| 1g: appearance | `.appearance.json` | non-empty | processor_results > 0 |
|
||||
| 1h: face_dedup | `.face_cluster.json` | non-empty | face_clusters > 0 |
|
||||
| 2a: face_trace | `.face_traced.json` | non-empty | Qdrant `_faces` trace_id set |
|
||||
| 2b: rule1 | `.rule1.json` | non-empty | chunk (sentence) > 0 |
|
||||
| 2c: vectorize | `.vectorize.json` | non-empty | chunk_vectors > 0 |
|
||||
| 2d: identity_agent | `.identity_agent.json` | non-empty | file_identities > 0 |
|
||||
| 3a: tkg_nodes | `.tkg_nodes.json` | non-empty | tkg_nodes > 0 |
|
||||
| 3b: tkg_edges | `.tkg_edges.json` | non-empty | tkg_edges > 0 |
|
||||
| 4a: rule2 | `.rule2.json` | non-empty | chunk (relationship) > 0 |
|
||||
|
||||
### 4.3 Status Values
|
||||
|
||||
| Status | Meaning |
|
||||
|--------|---------|
|
||||
| `pending` | Not yet started |
|
||||
| `running` | Currently executing |
|
||||
| `completed` | L1 + L2 + L3 all pass |
|
||||
| `failed` | L1 passes but L2 or L3 fails |
|
||||
| `missing` | L1 fails (artifact not on disk) |
|
||||
| `skipped` | Optional stage not run |
|
||||
|
||||
---
|
||||
|
||||
## 5. Rebuild
|
||||
|
||||
### 5.1 Rebuild Principle
|
||||
|
||||
Any sub-stage can be rebuilt independently:
|
||||
1. Read input artifacts (from disk or DB)
|
||||
2. Re-run the stage logic (processor or post-processor)
|
||||
3. Write output artifact + update DB/Qdrant
|
||||
|
||||
### 5.2 Rebuild Dependency
|
||||
|
||||
To rebuild stage N, all its dependency stages must be `completed`:
|
||||
|
||||
| Stage | Required Dependencies |
|
||||
|-------|----------------------|
|
||||
| 0a-0c | video file on disk |
|
||||
| 1a: cut | 0a (probe) |
|
||||
| 1b: asr | video file |
|
||||
| 1c: asrx | 1a (cut) + 1b (asr) |
|
||||
| 1d: ocr | video file |
|
||||
| 1e: face | video file |
|
||||
| 1f: pose | 1e (face) |
|
||||
| 1g: appearance | 1f (pose) |
|
||||
| 1h: face_dedup | 1e (face) |
|
||||
| 2a: face_trace | 1e (face) |
|
||||
| 2b: rule1 | 1c (asrx) |
|
||||
| 2c: vectorize | 2b (rule1) — chunks in DB |
|
||||
| 2d: identity_agent | 1h (face_dedup) — optional |
|
||||
| 3a: tkg_nodes | 1e (face), 2a (face_trace), 1c (asrx), 1d (ocr), 1g (appearance) |
|
||||
| 3b: tkg_edges | 3a (tkg_nodes) + 1c (asrx) |
|
||||
| 4a: rule2 | 3b (tkg_edges) + 2b (rule1) — chunks in DB |
|
||||
| 5a: complete | All required stages completed |
|
||||
|
||||
### 5.3 Rebuild API
|
||||
|
||||
```
|
||||
{OUTPUT_DIR}/
|
||||
{file_uuid}.probe.json ← ffprobe
|
||||
{file_uuid}.face.json ← face detection
|
||||
{file_uuid}.pre.json ← pre-processor (NEW)
|
||||
POST /api/v1/file/:file_uuid/rebuild/:stage
|
||||
```
|
||||
|
||||
- Validates dependencies are met
|
||||
- Re-runs the stage
|
||||
- Returns updated verification status
|
||||
|
||||
### 5.4 Rebuild via CLI
|
||||
|
||||
```bash
|
||||
# Check all stages
|
||||
python3 scripts/lifecycle_check.py --file-uuid <UUID>
|
||||
|
||||
# Rebuild specific stage
|
||||
python3 scripts/lifecycle_check.py --file-uuid <UUID> --rebuild 1c
|
||||
|
||||
# Rebuild from first missing stage
|
||||
python3 scripts/lifecycle_check.py --file-uuid <UUID> --rebuild auto
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Frontend Display
|
||||
|
||||
### 6.1 Two-Layer Architecture
|
||||
|
||||
**Layer 1: High-Level Summary** (default view)
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ ▶ S0: Register ████████████ 3/3 completed │
|
||||
│ ▶ S1: Processors ████████░░░░ 6/8 partial │
|
||||
│ ▶ S2: Post-Process ██░░░░░░░░░░ 1/4 running │
|
||||
│ ▶ S3: TKG Build ░░░░░░░░░░░░ 0/2 pending │
|
||||
│ ▶ S4: Rule2 ░░░░░░░░░░░░ 0/1 pending │
|
||||
│ ▶ S5: Complete ░░░░░░░░░░░░ 0/1 pending │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Layer 2: Expandable Sub-Stages** (click to expand)
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────┐
|
||||
│ ▼ S1: Processors ████████░░░░ 6/8 partial │
|
||||
│ ├─ 1a: CUT ✅ completed │
|
||||
│ ├─ 1b: ASR ✅ completed │
|
||||
│ ├─ 1c: ASRX ✅ completed │
|
||||
│ ├─ 1d: OCR ✅ completed │
|
||||
│ ├─ 1e: Face ✅ completed │
|
||||
│ ├─ 1f: Pose ✅ completed │
|
||||
│ ├─ 1g: Appearance ❌ missing │
|
||||
│ └─ 1h: Face Dedup ⏭ skipped (manual) │
|
||||
└─────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### 6.2 Sub-Stage Display Names
|
||||
|
||||
| Code Name | Display Name |
|
||||
|-----------|-------------|
|
||||
| probe | Probe (ffprobe) |
|
||||
| profile | File Profile |
|
||||
| key_frame | Key Frame |
|
||||
| cut | Scene Detection (CUT) |
|
||||
| asr | Speech Recognition (ASR) |
|
||||
| asrx | Speaker Diarization (ASRX) |
|
||||
| ocr | Text Recognition (OCR) |
|
||||
| face | Face Detection |
|
||||
| pose | Pose Estimation |
|
||||
| appearance | Appearance Features |
|
||||
| face_dedup | Face Deduplication |
|
||||
| face_trace | Face Tracking |
|
||||
| rule1 | Rule1 Ingestion |
|
||||
| vectorize | Vector Embedding |
|
||||
| identity_agent | Identity Agent |
|
||||
| tkg_nodes | TKG Nodes |
|
||||
| tkg_edges | TKG Edges |
|
||||
| rule2 | Rule2 Ingestion |
|
||||
| complete | Complete |
|
||||
|
||||
### 6.3 API Contract
|
||||
|
||||
The frontend fetches stage data from:
|
||||
|
||||
```
|
||||
GET /api/v1/stats/pipeline/:file_uuid
|
||||
```
|
||||
|
||||
Response:
|
||||
```json
|
||||
{
|
||||
"file_name": "charade.mp4",
|
||||
"file_path": "/data/demo/charade.mp4",
|
||||
"canonical_path": "/private/data/demo/charade.mp4",
|
||||
"content_hash": "a1b2c3d4e5f6...",
|
||||
"probe_json": {
|
||||
"format": { "duration": "6879.3", "size": "2147483648" },
|
||||
"streams": [...]
|
||||
},
|
||||
"birthday": "2026-05-15T02:15:00Z",
|
||||
"file_uuid": "aeed71342a899fe4b4c57b7d41bcb692",
|
||||
"file_size": 2147483648,
|
||||
"file_type": "video | image | document | audio",
|
||||
"pre_processed_at": "2026-05-15T02:15:05Z"
|
||||
"file_uuid": "abc123",
|
||||
"overall_progress": 0.45,
|
||||
"stages": [
|
||||
{
|
||||
"name": "register",
|
||||
"weight": 0.05,
|
||||
"progress": 1.0,
|
||||
"status": "completed",
|
||||
"detail": "3/3 sub-stages",
|
||||
"sub_stages": [
|
||||
{"name": "probe", "status": "completed", "artifact": "probe.json"},
|
||||
{"name": "profile", "status": "completed", "artifact": "profile.json"},
|
||||
{"name": "key_frame", "status": "completed", "artifact": "key_frame.jpg"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"name": "processors",
|
||||
"weight": 0.40,
|
||||
"progress": 0.75,
|
||||
"status": "partial",
|
||||
"detail": "6/8 sub-stages",
|
||||
"sub_stages": [
|
||||
{"name": "cut", "status": "completed", "artifact": "cut.json"},
|
||||
{"name": "asr", "status": "completed", "artifact": "asr.json"},
|
||||
{"name": "asrx", "status": "completed", "artifact": "asrx.json"},
|
||||
{"name": "ocr", "status": "completed", "artifact": "ocr.json"},
|
||||
{"name": "face", "status": "completed", "artifact": "face.json"},
|
||||
{"name": "pose", "status": "completed", "artifact": "pose.json"},
|
||||
{"name": "appearance", "status": "missing", "artifact": "appearance.json"},
|
||||
{"name": "face_dedup", "status": "skipped", "artifact": "face_cluster.json"}
|
||||
]
|
||||
}
|
||||
],
|
||||
"updated_at": "2026-07-22T18:00:00Z"
|
||||
}
|
||||
```
|
||||
|
||||
### Key Design: file_uuid = f(mac, birthday, path, filename)
|
||||
---
|
||||
|
||||
The `birthday` is `file modification time` (mtime) — obtained from `fs::metadata().modified()`. Using mtime instead of birthtime ensures file_uuid stability when files are transferred between systems via rsync (which preserves mtime but not birthtime on macOS).
|
||||
## 7. Weight Distribution
|
||||
|
||||
### 7.1 High-Level Stage Weights
|
||||
|
||||
| Stage | Weight | Rationale |
|
||||
|-------|--------|-----------|
|
||||
| S0: Register | 5% | Fast, prerequisite for everything |
|
||||
| S1: Processors | 40% | Most time-consuming, GPU-bound |
|
||||
| S2: Post-Process | 20% | Face trace + Rule1 + Vectorize |
|
||||
| S3: TKG Build | 20% | Node + edge construction |
|
||||
| S4: Rule2 | 10% | Relationship chunk creation |
|
||||
| S5: Complete | 5% | Final status update |
|
||||
|
||||
### 7.2 Processor Sub-Weights (within S1 = 40%)
|
||||
|
||||
| Processor | Sub-Weight | Rationale |
|
||||
|-----------|-----------|-----------|
|
||||
| CUT | 5% | Scene detection, ~10s |
|
||||
| ASR | 15% | whisper-small, ~2min/10min video |
|
||||
| ASRX | 20% | Speaker diarization, ~3min |
|
||||
| OCR | 10% | PaddleOCR, ~1min |
|
||||
| Face | 15% | CoreML FaceNet, ~1min |
|
||||
| Pose | 10% | mediapipe, ~1min |
|
||||
| Appearance | 5% | Feature extraction, ~30s |
|
||||
| Face Dedup | 0% | Manual (not in automated pipeline) |
|
||||
|
||||
---
|
||||
|
||||
## 8. Artifact Naming Convention
|
||||
|
||||
All artifacts live in the output directory (`MOMENTRY_OUTPUT_DIR`):
|
||||
|
||||
```
|
||||
birthday = 2026-05-15T02:15:00Z ← file birth time, never changes
|
||||
↓
|
||||
file_uuid = SHA256(mac | birthday | path | filename)
|
||||
↓
|
||||
Same file: same path + filename → same file_uuid, regardless of registration count
|
||||
Different files: different content_hash → different file_uuid (even if same name)
|
||||
{output_dir}/
|
||||
├─ {uuid}.probe.json # S0: ffprobe metadata
|
||||
├─ {uuid}.profile.json # S0: FileProfile
|
||||
├─ {uuid}.key_frame.jpg # S0: extracted key frame
|
||||
├─ {uuid}.cut.json # S1: scene boundaries
|
||||
├─ {uuid}.asr.json # S1: speech transcription
|
||||
├─ {uuid}.asrx.json # S1: speaker diarization
|
||||
├─ {uuid}.ocr.json # S1: text detections
|
||||
├─ {uuid}.face.json # S1: face detections + embeddings
|
||||
├─ {uuid}.face_cluster.json # S1: face clustering (optional)
|
||||
├─ {uuid}.pose.json # S1: pose estimations
|
||||
├─ {uuid}.appearance.json # S1: appearance features
|
||||
├─ {uuid}.face_traced.json # S2: face tracking with trace_id
|
||||
├─ {uuid}.rule1.json # S2: sentence chunks
|
||||
├─ {uuid}.vectorize.json # S2: embedding stats
|
||||
├─ {uuid}.identity_agent.json # S2: identity matching (optional)
|
||||
├─ {uuid}.tkg_nodes.json # S3: TKG node dump
|
||||
├─ {uuid}.tkg_edges.json # S3: TKG edge dump
|
||||
├─ {uuid}.rule2.json # S4: relationship chunks
|
||||
└─ {uuid}/ # Trace profiles directory
|
||||
├─ trace_0/
|
||||
│ ├─ trace_profile.json
|
||||
│ ├─ key_frame.jpg
|
||||
│ └─ key_face.jpg
|
||||
├─ trace_1/
|
||||
│ └─ ...
|
||||
└─ trace_N/
|
||||
```
|
||||
|
||||
## Phase 2: Registration (Citizenship)
|
||||
---
|
||||
|
||||
### POST /api/v1/files/register
|
||||
## 9. Current State Audit (Gamma 8)
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:3002/api/v1/files/register \
|
||||
-H "X-API-Key: ..." \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"file_path":"/data/demo/charade.mp4"}'
|
||||
```
|
||||
File: `d3f9ae8e471a1fc4d47022c66091b920` (Gamma 8-Director Chih-Lin Yang)
|
||||
|
||||
### Flow
|
||||
| Sub-Stage | Artifact | Status |
|
||||
|-----------|----------|--------|
|
||||
| 0a: probe | probe.json | ✅ exists |
|
||||
| 0b: profile | profile.json | ❌ missing |
|
||||
| 0c: key_frame | key_frame.jpg | ❌ missing |
|
||||
| 1a: cut | cut.json | ✅ exists |
|
||||
| 1b: asr | asr.json | ✅ exists |
|
||||
| 1c: asrx | asrx.json | ✅ exists |
|
||||
| 1d: ocr | ocr.json | ✅ exists |
|
||||
| 1e: face | face.json | ✅ exists |
|
||||
| 1f: pose | pose.json | ✅ exists |
|
||||
| 1g: appearance | appearance.json | ❌ missing |
|
||||
| 1h: face_dedup | face_cluster.json | ⏭ skipped (manual) |
|
||||
| 2a: face_trace | face_traced.json | ✅ exists |
|
||||
| 2b: rule1 | rule1.json | ❌ missing |
|
||||
| 2c: vectorize | vectorize.json | ❌ missing |
|
||||
| 2d: identity_agent | identity_agent.json | ❌ missing |
|
||||
| 3a: tkg_nodes | tkg_nodes.json | ❌ missing |
|
||||
| 3b: tkg_edges | tkg_edges.json | ❌ missing |
|
||||
| 4a: rule2 | rule2.json | ❌ missing |
|
||||
|
||||
```
|
||||
1. Check {OUTPUT_DIR}/{file_uuid}.pre.json
|
||||
├─ Exists AND content_hash matches → use cached (skip SHA256 + probe)
|
||||
└─ Not exists OR hash mismatch → compute fresh (existing logic)
|
||||
|
||||
2. Dedup check: SELECT file_uuid FROM videos WHERE content_hash = $1
|
||||
├─ Found → already_exists: true (identical DNA = same file)
|
||||
└─ Not found → continue
|
||||
|
||||
3. Name conflict check + auto-rename if needed
|
||||
└─ charade.mp4 → charade (1).mp4 (same name, different content)
|
||||
|
||||
4. INSERT INTO videos (
|
||||
file_uuid, file_path, file_name, file_type,
|
||||
duration, width, height, fps,
|
||||
probe_json, content_hash, status, registration_time
|
||||
) VALUES (
|
||||
$1, $2, $3, $4, $5, $6, $7, $8, $9, $10,
|
||||
'registered', NOW() ← status=registered, registration_time=NOW()
|
||||
)
|
||||
```
|
||||
|
||||
## Data Separation
|
||||
|
||||
| Field | Source | Computed When | Mutable |
|
||||
|-------|--------|---------------|:------:|
|
||||
| `birthday` | `fs::metadata().modified()` (mtime) | Pre-process (once) | ❌ Never (stable across rsync) |
|
||||
| `content_hash` (SHA256) | Full file | Pre-process (once) | ❌ Never (unless file modified) |
|
||||
| `file_uuid` | SHA256(mac\|birthday\|path\|filename) | Pre-process (once) | ❌ Never |
|
||||
| `registration_time` | `NOW()` at register | Register API | ✅ Per registration |
|
||||
| `status` | — | Register API | `unregistered` → `registered` |
|
||||
|
||||
## File Lifecycle State Diagram
|
||||
|
||||
```
|
||||
File detected by watcher (detection only, no modification)
|
||||
│
|
||||
│ Pre-processing triggered explicitly (API or register)
|
||||
▼
|
||||
[Pre-Processor]
|
||||
├─ SHA256 (DNA / fingerprint)
|
||||
├─ ffprobe (metadata extraction)
|
||||
└─ file_uuid (birth certificate ID)
|
||||
│
|
||||
▼
|
||||
{file_uuid}.pre.json
|
||||
status = unregistered (no DB record)
|
||||
│
|
||||
│ (user calls POST /api/v1/files/register)
|
||||
▼
|
||||
[Register Handler]
|
||||
├─ Read .pre.json → skip recomputation
|
||||
├─ Dedup check (content_hash collision?)
|
||||
├─ Name check + rename?
|
||||
└─ INSERT INTO videos
|
||||
│
|
||||
▼
|
||||
status = registered
|
||||
registration_time = NOW()
|
||||
```
|
||||
|
||||
## Implementation Checklist
|
||||
|
||||
| # | Task | File |
|
||||
|---|------|------|
|
||||
| 1 | Expose `pre_process_file()` as public function (SHA256 + probe + file_uuid → `.pre.json`) | `src/watcher/watcher.rs` |
|
||||
| 2 | Register: read `.pre.json`, skip SHA256/probe if cached | `src/api/server.rs` → `register_single_file` |
|
||||
| 3 | file_uuid: use `birthday` from `.pre.json` (or `fs::metadata().modified()` fallback) | `src/api/server.rs` |
|
||||
| 4 | INSERT status: `registered`, registration_time: `NOW()` | `src/api/server.rs` |
|
||||
**Observations**:
|
||||
- S1 processors mostly complete, but Appearance missing (1g)
|
||||
- S0 profile/key_frame missing (registration may not have created them)
|
||||
- S2-S4 all have DB data but no flat `.json` dumps
|
||||
|
||||
---
|
||||
|
||||
## Version History
|
||||
|
||||
| Version | Date | Changes |
|
||||
|---------|------|---------|
|
||||
| V1.0 | 2026-05-15 | Initial design — birth certificate (pre-process) + civil registration two-phase flow |
|
||||
| V1.1 | 2026-05-15 | Reclassified from DESIGN to STANDARDS as design standard |
|
||||
| V1.2 | 2026-05-15 | mtime replaces birthtime for file_uuid stability across rsync; watcher is detection-only |
|
||||
| Version | Date | Author | Changes |
|
||||
|---------|------|--------|---------|
|
||||
| 1.2 | 2026-07-22 | OpenCode | Added CUT scene key frames extraction for VLM analysis |
|
||||
| 1.1 | 2026-07-22 | OpenCode | Added S0b: audio_track classification (VAD) — 6 stages, 15 sub-stages |
|
||||
| 1.0 | 2026-07-22 | OpenCode | Initial design — 6 stages, 14 sub-stages, I/O specs, verification, rebuild |
|
||||
|
||||
@@ -0,0 +1,153 @@
|
||||
# File Profile V1.0
|
||||
|
||||
**Status:** Active
|
||||
**Version:** 1.0
|
||||
**Date:** 2026-07-22
|
||||
**Scope:** File identity artifact — persistent on-disk profile per registered file
|
||||
|
||||
---
|
||||
|
||||
## Problem
|
||||
|
||||
- 5 zombie files in DB: `file_uuid` exists but `file_name` and `file_path` are empty — no way to recover
|
||||
- File identity lives only in PostgreSQL; no on-disk fallback
|
||||
- `birth_registration` written by `ingestion.rs` but **not** by `files.rs` API path
|
||||
- No file history — if a file moves or is renamed, no record of where it was
|
||||
|
||||
## Design
|
||||
|
||||
A JSON file created **first** during registration, stored flat in `MOMENTRY_OUTPUT_DIR`:
|
||||
|
||||
```
|
||||
{MOMENTRY_OUTPUT_DIR}/{file_uuid}.profile.json
|
||||
```
|
||||
|
||||
### JSON Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "1.0",
|
||||
"file_uuid": "84d838f260e1881a0daa55fabbc8e434",
|
||||
"file_name": "view28.mp4",
|
||||
"file_type": "video",
|
||||
"birth": {
|
||||
"mac_address": "a1:b2:c3:d4:e5:f6",
|
||||
"birthday": "2026-04-13T23:00:49+08:00",
|
||||
"original_path": "/Users/accusys/momentry/var/sftpgo/data/demo",
|
||||
"original_filename": "view28.mp4",
|
||||
"canonical_path": "/Users/accusys/momentry/var/sftpgo/data/demo/view28.mp4",
|
||||
"content_hash": "abc123..."
|
||||
},
|
||||
"current": {
|
||||
"path": "/Users/accusys/momentry/var/sftpgo/data/demo/view28.mp4",
|
||||
"file_name": "view28.mp4",
|
||||
"file_type": "video"
|
||||
},
|
||||
"history": [
|
||||
{
|
||||
"action": "registered",
|
||||
"timestamp": "2026-07-22T14:30:00+08:00",
|
||||
"path": "/Users/accusys/momentry/var/sftpgo/data/demo/view28.mp4",
|
||||
"file_name": "view28.mp4"
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"duration": 243.24,
|
||||
"width": 720,
|
||||
"height": 890,
|
||||
"fps": 60.0,
|
||||
"total_frames": 7297
|
||||
},
|
||||
"key_frame": null
|
||||
}
|
||||
```
|
||||
|
||||
### Fields
|
||||
|
||||
| Field | Purpose |
|
||||
|-------|---------|
|
||||
| `version` | Profile schema version (for future migration) |
|
||||
| `file_uuid` | The deterministic UUID |
|
||||
| `file_name` | Original filename at registration |
|
||||
| `file_type` | video/audio/image/document/... |
|
||||
| `birth.mac_address` | MAC address used to compute UUID |
|
||||
| `birth.birthday` | File mtime (RFC3339) used to compute UUID |
|
||||
| `birth.original_path` | Parent directory at registration |
|
||||
| `birth.original_filename` | Filename at registration |
|
||||
| `birth.canonical_path` | Canonical (resolved symlinks) path at registration |
|
||||
| `birth.content_hash` | SHA256 of file content |
|
||||
| `current.path` | Latest known path (updated when file moves) |
|
||||
| `current.file_name` | Latest known filename (updated on rename) |
|
||||
| `current.file_type` | Latest file type |
|
||||
| `history` | Array of all path/name changes with timestamps |
|
||||
| `metadata` | Media info (duration, resolution, etc.) |
|
||||
| `key_frame` | Base64-encoded JPEG of representative frame (video only), or null |
|
||||
|
||||
### key_frame
|
||||
|
||||
For video files, `key_frame` stores a **base64-encoded JPEG** of a representative frame extracted at registration time (typically at 10% of duration or the first non-black frame). For non-video files, this field is `null`.
|
||||
|
||||
Purpose:
|
||||
- Instant visual identification without needing to open the video
|
||||
- Fallback if thumbnails or `.faces/` crops are deleted
|
||||
- Portable — the profile file is self-contained
|
||||
|
||||
Extraction:
|
||||
- Uses ffmpeg to grab a frame at `duration * 0.1` (or first frame if duration unknown)
|
||||
- JPEG quality 85, max width 640px
|
||||
- Stored inline as base64 string in the JSON
|
||||
|
||||
## Implementation
|
||||
|
||||
### New Module
|
||||
|
||||
`src/core/file_profile.rs` — `FileProfile` struct with:
|
||||
- `from_registration_params(...)` — build at registration time
|
||||
- `load_from_disk(uuid, output_dir)` — read `{uuid}.profile.json`
|
||||
- `save_to_disk(&self, output_dir)` — write `{uuid}.profile.json`
|
||||
- `update_current_path(&mut self, new_path, new_name)` — append to history, update current
|
||||
- `extract_key_frame(video_path, duration)` — ffmpeg frame extraction + base64
|
||||
|
||||
### Registration Flow
|
||||
|
||||
1. DB INSERT (existing)
|
||||
2. **Build FileProfile** from all params (mac, birthday, path, name, content_hash, probe metadata)
|
||||
3. **Extract key_frame** if video (ffmpeg)
|
||||
4. **Save `{uuid}.profile.json`** — first artifact on disk
|
||||
5. CUT processing (existing)
|
||||
|
||||
### Update Flow
|
||||
|
||||
When `file_path` or `file_name` changes via API:
|
||||
1. Load profile from disk
|
||||
2. `profile.update_current_path(new_path, new_name)`
|
||||
3. Save updated profile
|
||||
|
||||
### Fallback Flow
|
||||
|
||||
If DB data is missing (zombie files):
|
||||
1. Load profile from disk
|
||||
2. Profile's `current.path` and `current.file_name` provide recovery data
|
||||
|
||||
### Files Changed
|
||||
|
||||
| File | Change |
|
||||
|------|--------|
|
||||
| `src/core/file_profile.rs` | **NEW** — FileProfile struct |
|
||||
| `src/core/mod.rs` | Add `pub mod file_profile` |
|
||||
| `src/api/files.rs` | Write profile after registration; fallback; enrich GET; cleanup |
|
||||
| `src/core/ingestion.rs` | Write profile after registration |
|
||||
| `src/api/profile.rs` | Enrich GET file-profile with profile data + history |
|
||||
|
||||
### Backfill
|
||||
|
||||
Existing 16 files get profiles generated from DB fields + probe.json data.
|
||||
5 zombie files get minimal profiles (UUID + file_type + content_hash from DB).
|
||||
|
||||
---
|
||||
|
||||
## Version History
|
||||
|
||||
| Version | Date | Change |
|
||||
|---------|------|--------|
|
||||
| 1.0 | 2026-07-22 | Initial design — file profile with key_frame |
|
||||
@@ -0,0 +1,339 @@
|
||||
# Face-Pose-Appearance Tracking Design
|
||||
|
||||
**Version**: 1.0
|
||||
**Date**: 2026-07-19
|
||||
**Status**: Ready for Implementation
|
||||
|
||||
---
|
||||
|
||||
## 1. Overview
|
||||
|
||||
本文件定義 Face、Pose、Appearance 的追蹤系統設計,包含:
|
||||
- Trace ID 繼承規則
|
||||
- 擴張邏輯
|
||||
- Appearance 色塊提取
|
||||
- Agent Search 整合
|
||||
|
||||
---
|
||||
|
||||
## 2. Core Concepts
|
||||
|
||||
### 2.1 Identity vs Tracking
|
||||
|
||||
| Processor | Purpose | Description |
|
||||
|-----------|---------|-------------|
|
||||
| **Face** | Identity | Who is this person? 需要高品質 embedding |
|
||||
| **Pose** | Tracking | Where is this person? 當 face occluded 時維持追蹤 |
|
||||
| **Appearance** | Tracking | What do they look like? 當 pose occluded 時維持追蹤 |
|
||||
|
||||
### 2.2 Offline Processing Advantage
|
||||
|
||||
Offline 處理可以先做 face detection,再從 face traces 擴張 pose/appearance:
|
||||
|
||||
```
|
||||
Face Detection → 知道身份錨點
|
||||
↓
|
||||
Face Tracking → 給予 trace_id
|
||||
↓
|
||||
Pose Expansion → 從 face traces 向外擴張
|
||||
↓
|
||||
Appearance Expansion → 從 pose traces 向外擴張
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Processing Pipeline
|
||||
|
||||
### 3.1 Pipeline Order
|
||||
|
||||
| Order | Processor | Dependencies | Output | Description |
|
||||
|-------|-----------|-------------|--------|-------------|
|
||||
| 1 | `cut` | — | cut.json | Scene detection |
|
||||
| 2 | `face` | — | face.json | Face detection (8Hz) + embedding |
|
||||
| 3 | `face_trace` | face | face_traced.json | Face tracking (IoU + embedding) |
|
||||
| 4 | `pose` | face_trace | pose.json | Pose expansion from traces |
|
||||
| 5 | `appearance` | pose | appearance.json | Appearance extraction |
|
||||
| 6 | `asr` | cut | asr.json | Speech-to-text |
|
||||
| 7 | `asrx` | asr | asrx.json | Speaker diarization |
|
||||
|
||||
### 3.2 Sampling Rate
|
||||
|
||||
- **公式**: `sample_interval = floor(fps / 8)`
|
||||
- **確保**: ≥ 8Hz 取樣率
|
||||
- **範例**:
|
||||
- 24fps → interval = 3 → 8Hz
|
||||
- 30fps → interval = 3 → 10Hz
|
||||
- 60fps → interval = 7 → 8.6Hz
|
||||
|
||||
---
|
||||
|
||||
## 4. Trace ID Inheritance
|
||||
|
||||
### 4.1 Inheritance Chain
|
||||
|
||||
```
|
||||
Face Trace (identity anchor)
|
||||
│ trace_id = 1, 2, 3, ...
|
||||
│
|
||||
▼ inherits trace_id
|
||||
Pose Expansion
|
||||
│ same trace_id per person
|
||||
│
|
||||
▼ inherits trace_id
|
||||
Appearance Expansion
|
||||
│ same trace_id per person
|
||||
```
|
||||
|
||||
### 4.2 Frame Count Relationship
|
||||
|
||||
```
|
||||
face_frames ≤ pose_frames ≤ appearance_frames
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- Face: 只有臉部可見的 frames
|
||||
- Pose: Face frames + 擴張 frames(臉被遮但身體可見)
|
||||
- Appearance: Pose frames + 擴張 frames
|
||||
|
||||
### 4.3 Trace Connection
|
||||
|
||||
```
|
||||
Face trace A (frames 1-10) Face trace B (frames 20-30)
|
||||
↘ ↙
|
||||
Pose 連接 (frames 15-18)
|
||||
(同一人,中間臉被遮住)
|
||||
```
|
||||
|
||||
**意義**: Pose 可以連接斷開的 face traces,屬於同一人。
|
||||
|
||||
---
|
||||
|
||||
## 5. Expansion Rules
|
||||
|
||||
### 5.1 Pose Expansion
|
||||
|
||||
**Algorithm**:
|
||||
1. 讀取 face_traced.json,取得每個 trace_id 的 frames
|
||||
2. 對每個 trace 的 frames 向外擴張(逐幀檢查)
|
||||
3. 連續 3 幀無 pose detection → 停止擴張
|
||||
4. 繼承 trace_id
|
||||
5. 輸出 8Hz 取樣
|
||||
|
||||
**Parameters**:
|
||||
|
||||
| Parameter | Value | Description |
|
||||
|-----------|-------|-------------|
|
||||
| `miss_threshold` | 3 | 連續無檢測幀數 |
|
||||
| `output_rate` | 8Hz | 輸出取樣率 |
|
||||
|
||||
### 5.2 Appearance Expansion
|
||||
|
||||
**Algorithm**:
|
||||
1. 讀取 pose.json,取得每個 trace_id 的 frames
|
||||
2. 對每個 pose frame,在 keypoint 位置提取顏色
|
||||
3. 記錄整體亮度
|
||||
4. 輸出 8Hz 取樣
|
||||
|
||||
**Parameters**:
|
||||
|
||||
| Parameter | Value | Description |
|
||||
|-----------|-------|-------------|
|
||||
| `color_radius` | 15 | 顏色取樣半徑(pixels) |
|
||||
| `output_rate` | 8Hz | 輸出取樣率 |
|
||||
|
||||
---
|
||||
|
||||
## 6. Pose Output
|
||||
|
||||
### 6.1 Data Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"frame_count": 1000,
|
||||
"fps": 24.0,
|
||||
"frames": [
|
||||
{
|
||||
"frame": 100,
|
||||
"timestamp": 4.16,
|
||||
"trace_id": 1,
|
||||
"persons": [
|
||||
{
|
||||
"keypoints": [
|
||||
{"name": "nose", "x": 315.9, "y": 364.2, "confidence": 0.85},
|
||||
{"name": "left_shoulder", "x": 290.0, "y": 400.0, "confidence": 0.92}
|
||||
],
|
||||
"bbox": {"x": 280, "y": 350, "width": 100, "height": 200}
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 6.2 Bbox Validation
|
||||
|
||||
**原則**: Face bbox 應在 Pose bbox 內,或 IoU > 0.5
|
||||
|
||||
```
|
||||
┌─────────────────────────┐
|
||||
│ Pose bbox │
|
||||
│ ┌─────────┐ │
|
||||
│ │ Face │ │
|
||||
│ │ bbox │ │
|
||||
│ └─────────┘ │
|
||||
└─────────────────────────┘
|
||||
```
|
||||
|
||||
**用途**:
|
||||
- 品質驗證:確保 pose/face 屬於同一人
|
||||
- 匹配追蹤:用 bbox overlap 匹配 face/pose
|
||||
|
||||
---
|
||||
|
||||
## 7. Appearance Output
|
||||
|
||||
### 7.1 Keypoint-based Color Extraction
|
||||
|
||||
**原理**: 在 pose keypoint 位置取周圍平均色
|
||||
|
||||
```
|
||||
Pose Keypoints 座標
|
||||
↓
|
||||
在每個 keypoint 位置取色
|
||||
↓
|
||||
記錄為 appearance
|
||||
```
|
||||
|
||||
### 7.2 Body Part Mapping
|
||||
|
||||
| Keypoints | Body Part | Description |
|
||||
|-----------|-----------|-------------|
|
||||
| nose, eyes, ears | head | 帽子、頭髮顏色 |
|
||||
| shoulders | torso | 上衣顏色 |
|
||||
| hips, knees | legs | 褲子顏色 |
|
||||
| ankles | feet | 鞋子顏色 |
|
||||
|
||||
### 7.3 Data Structure
|
||||
|
||||
```json
|
||||
{
|
||||
"frame_count": 1000,
|
||||
"fps": 24.0,
|
||||
"frames": [
|
||||
{
|
||||
"frame": 100,
|
||||
"timestamp": 4.16,
|
||||
"trace_id": 1,
|
||||
"brightness": 0.75,
|
||||
"colors": {
|
||||
"head": [180, 150, 120],
|
||||
"torso": [255, 50, 50],
|
||||
"legs": [50, 50, 200],
|
||||
"feet": [50, 200, 50]
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 7.4 Lighting Record
|
||||
|
||||
```json
|
||||
{
|
||||
"brightness": 0.75
|
||||
}
|
||||
```
|
||||
|
||||
**用途**: 不同光源下的顏色校正
|
||||
|
||||
---
|
||||
|
||||
## 8. VLM Complementary Strategy
|
||||
|
||||
### 8.1 Two-Level Approach
|
||||
|
||||
| Level | Method | Purpose |
|
||||
|-------|--------|---------|
|
||||
| **L1** | Keypoint 快取色 | 快速搜尋、初步候選 |
|
||||
| **L2** | VLM 驗證 | 複雜情況、細節補充(可選) |
|
||||
|
||||
### 8.2 Workflow
|
||||
|
||||
```
|
||||
搜尋「穿紅上衣的人」
|
||||
↓
|
||||
L1: Keypoint 取色搜尋 → Top 20 候選
|
||||
↓
|
||||
L2: VLM 驗證(需要時)→ 確認顏色、補充細節
|
||||
↓
|
||||
最終結果 → Top 10 + 置信度
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. Agent Search Integration
|
||||
|
||||
### 9.1 Agent Tool Design
|
||||
|
||||
```python
|
||||
def search_by_appearance(
|
||||
color: str, # "red", "blue", "green"
|
||||
body_part: str, # "torso", "legs", "feet"
|
||||
top_k: int = 10
|
||||
) -> List[SearchResult]:
|
||||
"""
|
||||
搜尋穿特定顏色衣物的人
|
||||
|
||||
Returns:
|
||||
[
|
||||
{"trace_id": 1, "identity": "John", "confidence": 0.85},
|
||||
{"trace_id": 2, "identity": "Mary", "confidence": 0.72},
|
||||
]
|
||||
"""
|
||||
```
|
||||
|
||||
### 9.2 Query Examples
|
||||
|
||||
| User Query | Agent Action |
|
||||
|------------|--------------|
|
||||
| 「穿紅上衣的人是誰?」 | search_by_appearance("red", "torso") → match identity |
|
||||
| 「穿綠鞋子的人」 | search_by_appearance("green", "feet") |
|
||||
| 「戴黑帽子的人」 | search_by_appearance("black", "head") |
|
||||
|
||||
### 9.3 Top-K Strategy
|
||||
|
||||
- **原則**: 找 top 10-20 最相似的
|
||||
- **容許誤差**: 光源、角度差異可接受
|
||||
- **近似即可**: 不需精確匹配
|
||||
|
||||
---
|
||||
|
||||
## 10. Implementation Files
|
||||
|
||||
| Component | File | Status |
|
||||
|-----------|------|--------|
|
||||
| Face Detection | `swift_face.swift` | ✅ Complete |
|
||||
| Face Tracking | `store_traced_faces.py` | ✅ Complete |
|
||||
| Pose Expansion | `swift_pose_expansion.swift` | ✅ Complete |
|
||||
| Appearance Expansion | `swift_appearance_expansion.swift` | ✅ Complete |
|
||||
| Pose Processor | `pose_processor_v2.py` | ✅ Complete |
|
||||
| Appearance Processor | `appearance_processor_v2.py` | ✅ Complete |
|
||||
|
||||
---
|
||||
|
||||
## 11. Testing Checklist
|
||||
|
||||
- [ ] 清除測試檔案重新註冊
|
||||
- [ ] 執行完整流程:face → trace → pose → appearance
|
||||
- [ ] 驗證 trace_id 繼承正確性
|
||||
- [ ] 驗證 frame count 關係 (face ≤ pose ≤ appearance)
|
||||
- [ ] 驗證 bbox 包含關係 (face bbox ⊂ pose bbox)
|
||||
- [ ] 測試 Agent search_by_appearance
|
||||
|
||||
---
|
||||
|
||||
## 12. Version History
|
||||
|
||||
| Version | Date | Changes |
|
||||
|---------|------|---------|
|
||||
| 1.0 | 2026-07-19 | Initial design |
|
||||
@@ -0,0 +1,209 @@
|
||||
# Trace ID Inheritance & Expansion Rules
|
||||
|
||||
**Date**: 2026-07-19
|
||||
**Author**: Core Team
|
||||
**Status**: Final
|
||||
|
||||
---
|
||||
|
||||
## Overview
|
||||
|
||||
This document defines the trace ID inheritance rules and expansion logic for Face, Pose, and Appearance processing.
|
||||
|
||||
---
|
||||
|
||||
## Core Concepts
|
||||
|
||||
### Face = Identity, Pose/Appearance = Tracking
|
||||
|
||||
| Processor | Purpose | Description |
|
||||
|-----------|---------|-------------|
|
||||
| **Face** | Identity | Who is this person? Requires high-quality embedding for recognition. |
|
||||
| **Pose** | Tracking | Where is this person? Maintains tracking when face is occluded. |
|
||||
| **Appearance** | Tracking | What do they look like? Maintains tracking when pose is occluded. |
|
||||
|
||||
### Offline Processing Advantage
|
||||
|
||||
In offline processing, we can:
|
||||
1. First detect all faces (identity anchors)
|
||||
2. Then expand pose/appearance from face traces
|
||||
|
||||
This is different from real-time tracking where pose/appearance runs continuously and face anchors identity when visible.
|
||||
|
||||
---
|
||||
|
||||
## Processing Pipeline
|
||||
|
||||
### Step 1: Face Detection (8Hz)
|
||||
|
||||
```
|
||||
swift_face → face.json
|
||||
```
|
||||
|
||||
- Sampling rate: `floor(fps / 8)` (ensures ≥ 8Hz)
|
||||
- Output: Face bounding boxes with landmarks and embeddings
|
||||
|
||||
### Step 2: Face Tracking
|
||||
|
||||
```
|
||||
store_traced_faces.py → face_traced.json
|
||||
```
|
||||
|
||||
- Algorithm: IoU + embedding similarity
|
||||
- Output: Each face assigned a `trace_id`
|
||||
- Purpose: Group same-person faces across frames
|
||||
|
||||
### Step 3: Pose Expansion
|
||||
|
||||
```
|
||||
swift_pose_expansion → pose.json
|
||||
```
|
||||
|
||||
**Input**: face_traced.json (frames with trace_id)
|
||||
|
||||
**Expansion Algorithm**:
|
||||
1. For each trace_id, get all face frames
|
||||
2. Expand outward (forward/backward) checking for pose
|
||||
3. Stop when 3 consecutive frames have no pose detection
|
||||
4. Inherit trace_id from face
|
||||
|
||||
**Output**: Pose keypoints with inherited trace_id
|
||||
|
||||
### Step 4: Appearance Expansion
|
||||
|
||||
```
|
||||
swift_appearance_expansion → appearance.json
|
||||
```
|
||||
|
||||
**Input**: pose.json (frames with trace_id)
|
||||
|
||||
**Expansion Algorithm**:
|
||||
1. For each trace_id, get all pose frames
|
||||
2. Expand outward (forward/backward) checking for appearance
|
||||
3. Stop when 3 consecutive frames have HSV similarity < 0.5
|
||||
4. Inherit trace_id from pose
|
||||
|
||||
**Output**: HSV histograms with inherited trace_id
|
||||
|
||||
---
|
||||
|
||||
## Trace ID Inheritance
|
||||
|
||||
```
|
||||
Face Trace (identity anchor)
|
||||
│
|
||||
│ inherits trace_id
|
||||
▼
|
||||
Pose Expansion
|
||||
│
|
||||
│ inherits trace_id
|
||||
▼
|
||||
Appearance Expansion
|
||||
```
|
||||
|
||||
**Key Points:**
|
||||
- Trace ID originates from face tracking
|
||||
- Pose inherits the same trace_id (same person)
|
||||
- Appearance inherits the same trace_id (same person)
|
||||
- This enables linking all detections to the same identity
|
||||
|
||||
---
|
||||
|
||||
## Frame Count Relationship
|
||||
|
||||
```
|
||||
face frames ≤ pose frames ≤ appearance frames
|
||||
```
|
||||
|
||||
**Explanation:**
|
||||
- Face: Only frames where face is clearly visible
|
||||
- Pose: Face frames + expanded frames (pose may still be visible when face is occluded)
|
||||
- Appearance: Pose frames + expanded frames (appearance may still be visible when pose is occluded)
|
||||
|
||||
---
|
||||
|
||||
## Expansion Rules
|
||||
|
||||
### Pose Expansion
|
||||
|
||||
| Parameter | Value | Description |
|
||||
|-----------|-------|-------------|
|
||||
| `miss_threshold` | 3 | Stop after 3 consecutive frames without pose |
|
||||
| `max_range` | 300 frames | Maximum expansion distance (≈10s at 30fps) |
|
||||
| `output_rate` | 8Hz | Output sampling rate |
|
||||
|
||||
### Appearance Expansion
|
||||
|
||||
| Parameter | Value | Description |
|
||||
|-----------|-------|-------------|
|
||||
| `miss_threshold` | 3 | Stop after 3 consecutive frames with similarity < 0.5 |
|
||||
| `similarity_threshold` | 0.5 | HSV histogram similarity threshold |
|
||||
| `max_range` | 300 frames | Maximum expansion distance |
|
||||
| `output_rate` | 8Hz | Output sampling rate |
|
||||
|
||||
---
|
||||
|
||||
## Tracking Continuity
|
||||
|
||||
### Pose Can Connect Face Traces
|
||||
|
||||
```
|
||||
Face trace A (frames 1-10) Face trace B (frames 20-30)
|
||||
↘ ↙
|
||||
Pose connects (frames 15-18)
|
||||
(Same person, face was occluded)
|
||||
```
|
||||
|
||||
When pose expansion from two face traces overlaps, they may belong to the same person. Future enhancement: pose-based trace merging.
|
||||
|
||||
### Appearance Can Connect Pose Traces
|
||||
|
||||
Similar to pose, appearance similarity can connect pose traces when pose is temporarily occluded.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Files
|
||||
|
||||
| Component | File |
|
||||
|-----------|------|
|
||||
| Face Detection | `scripts/swift_processors/swift_face.swift` |
|
||||
| Face Tracking | `scripts/store_traced_faces.py` |
|
||||
| Pose Expansion | `scripts/swift_processors/swift_pose_expansion.swift` |
|
||||
| Appearance Expansion | `scripts/swift_processors/swift_appearance_expansion.swift` |
|
||||
| Pose Processor Wrapper | `scripts/pose_processor_v2.py` |
|
||||
| Appearance Processor Wrapper | `scripts/appearance_processor_v2.py` |
|
||||
| Dependencies Definition | `src/core/db/postgres_db.rs:568-577` |
|
||||
|
||||
---
|
||||
|
||||
## Testing
|
||||
|
||||
### Verify Trace ID Inheritance
|
||||
|
||||
```bash
|
||||
# Check face traces
|
||||
cat /Users/accusys/momentry/output/$FILE_UUID.face_traced.json | jq '.frames[].faces[].trace_id' | sort | uniq -c
|
||||
|
||||
# Check pose traces (should have same trace_ids)
|
||||
cat /Users/accusys/momentry/output/$FILE_UUID.pose.json | jq '.frames[].trace_id' | sort | uniq -c
|
||||
|
||||
# Check appearance traces (should have same trace_ids)
|
||||
cat /Users/accusys/momentry/output/$FILE_UUID.appearance.json | jq '.frames[].trace_id' | sort | uniq -c
|
||||
```
|
||||
|
||||
### Verify Frame Count Relationship
|
||||
|
||||
```bash
|
||||
# face ≤ pose ≤ appearance
|
||||
FACE_COUNT=$(cat $OUTPUT/$UUID.face.json | jq '.frames | length')
|
||||
POSE_COUNT=$(cat $OUTPUT/$UUID.pose.json | jq '.frames | length')
|
||||
APP_COUNT=$(cat $OUTPUT/$UUID.appearance.json | jq '.frames | length')
|
||||
|
||||
echo "Face: $FACE_COUNT, Pose: $POSE_COUNT, Appearance: $APP_COUNT"
|
||||
# Expected: Face ≤ Pose ≤ Appearance
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
*Document Version: 1.0*
|
||||
*Last Updated: 2026-07-19*
|
||||
Reference in New Issue
Block a user