fix: face group name read consistency, sync_file_status fix, cleanup ghost records, identity_agent replaced with face_dedup

- get_face_groups_handler: COALESCE(tp.name, tn.label) for name consistency
- sync_file_status: compare JSON vs pre_chunks (not chunk table)
- face consistency: compare frames.len() not total_faces
- cleanup 2 ghost records with NULL file_name/file_path
- replace identity_agent with face_dedup in pipeline stages
- remove identity_agent_api.rs and all references
- update required_processors to match actual processors
- update AGENTS.md with team responsibilities
- add Studio pipeline changes documentation
This commit is contained in:
Accusys
2026-07-27 02:15:51 +08:00
parent fcdeab82e6
commit 39a2cbc65b
118 changed files with 19386 additions and 2964 deletions
@@ -0,0 +1,148 @@
---
title: Always-Produce Processing Contract
version: 1.0
date: 2026-07-24
author: OpenCode
status: approved
---
# Always-Produce Processing Contract
## Scope
| Field | Value |
|-------|-------|
| Scope | All frame-based processors (face, pose, appearance, face_cluster, face_trace, etc.) |
| Status | Approved |
| Applies to | Python processors + Rust Worker |
| Related docs | `DESIGN/Processor_Module_V1.0.md`, `DESIGN/Redis_Progress_Reporting_V1.0.md`, `DESIGN/Worker_Health_Check_Mechanism.md` |
## 1. Frame-Scan Model
Video processing is fundamentally frame-based: a processor scans from frame 0 to the last frame.
```
Scan start → frame 0 → frame 1 → ... → frame N → scan complete
↓ ↓ ↓ ↓
Redis Redis Redis {uuid}.{p}.json
progress progress progress (final record)
```
### Key Rules
1. **Progress** = which frame has been scanned so far (`current_frame / total_frames`)
2. **Complete** = scanned to the last frame (proved by `.json` existing)
3. **Result** = always written, even if 0 detections found
## 2. Always-Produce Rule
### Principle
> Every processor MUST write its `{uuid}.{processor}.json` output file after completing its scan, **regardless of whether any results were found**.
### Rationale
The `.json` file serves dual purpose:
- **Proof of completion**: Worker uses `output_path.exists()` (line 580 of `job_worker.rs`) to skip already-finished processors
- **Downstream dependency**: Subsequent processors check this file for input
Without the Always-Produce rule:
- Zero-result processors leave no `.json` → Worker retries infinitely → deadlock
- Stuck jobs block downstream stages (Rule 1/2/3 ingestion, TKG build)
### Format
All processor JSON outputs MUST include:
```json
{
"status": "has_faces" | "no_faces" | "no_face_json" | "no_embeddings" | "success" | "error_*",
"file_uuid": "<uuid>",
...processor-specific fields (empty arrays when zero results)
}
```
Example — face cluster with no faces:
```json
{
"status": "no_faces",
"file_uuid": "9781de6d...",
"clusters": [],
"frames": []
}
```
### Processor Checklist
| Processor | Always-Produce? | Status field on 0 result |
|-----------|----------------|--------------------------|
| `face.py` | ✅ Yes | `"no_faces"` |
| `store_traced_faces.py` | ✅ Yes (already writes) | `"no_faces"` |
| `fast_face_clustering_processor.py` | ❌ **FIX NEEDED** | Early returns, no file written |
| `pose_processor*.py` | ✅ Yes | `"no_faces"` |
| `appearance_processor*.py` | ✅ Yes | `"no_faces"` |
## 3. Redis Progress During Scan
### Purpose
Live frame progress is published to Redis so the QC modal can display real-time status ("scanning frame 1234/5678").
### Mechanism
Use `redis_publisher.py` (`RedisPublisher` class) which publishes to Redis channel `{prefix}progress:{uuid}`:
```python
from redis_publisher import RedisPublisher
pub = RedisPublisher(file_uuid)
# During scan, per batch:
pub.progress("face_cluster", current_frame, total_frames, f"Scanning frame {current_frame}")
# On completion:
pub.complete("face_cluster", f"Done: {cluster_count} clusters")
```
### Frequency
- **Frame-based processors**: publish every N frames (batch/buffer flush)
- **Non-frame processors** (e.g., clustering): publish at meaningful milestones
## 4. Worker Heartbeat
### Problem
`health.rs` currently uses `check_process_running("worker")` which relies on `ps aux | grep momentry.*worker`. This is unreliable:
- Zombie processes show as "running"
- Stale matches from unrelated processes
### Fix
Worker writes a Redis HMSET `{prefix}health` with EXPIRE = `3 × poll_interval_secs` (default: 15s) in every `poll_and_process()` cycle.
Health endpoint checks:
1. Redis key `{prefix}health` exists
2. Key has remaining TTL > 0
3. Key's `status` field is `"healthy"` or `"throttled"`
If Redis key missing or expired → `worker_alive: false`.
## 5. Implementation Plan
| Step | File | Change |
|------|------|--------|
| 1 | `fast_face_clustering_processor.py` | Always-Produce for 3 early returns + Redis progress |
| 2 | `store_traced_faces.py` | Add Redis progress (optional) |
| 3 | `job_worker.rs` | Add EXPIRE after health HMSET |
| 4 | `health.rs` | Replace `check_process_running("worker")` with Redis TTL check |
| 5 | `processing.rs` | (Optional) Reject trigger if Worker not alive |
---
## Version History
| Version | Date | Author | Changes |
|---------|------|--------|---------|
| 1.0 | 2026-07-24 | OpenCode | Initial specification |
@@ -0,0 +1,341 @@
---
title: Face Tracking Pipeline Structure
version: 1.0
date: 2026-07-22
author: OpenCode
status: Active
---
# Face Tracking Pipeline — Structure Design
## Overview
```
Video
│
▼
┌──────────────────────────────┐
│ Stage 1: Face Detection │ face_processor.py
│ swift_face (Apple Vision) │ → {uuid}.face.json
│ CoreML FaceNet embedding │ → Qdrant _faces (initial)
└──────────────────────────────┘
│
▼
┌──────────────────────────────┐
│ Stage 2: Face Tracking │ store_traced_faces.py
│ face_tracker.py (IoU) │ → {uuid}.face_traced.json
│ trace_id assignment │ → Qdrant _faces (trace_id update)
└──────────────────────────────┘
│
▼
┌──────────────────────────────┐
│ Stage 3: Trace Profile │ backfill_trace_profiles.py
│ Qdrant _faces 分組 │ → output/{uuid}/trace_{N}/
│ key_frame + key_face │ trace_profile.json
└──────────────────────────────┘
│
▼
┌──────────────────────────────┐
│ Stage 4: TKG Nodes │ tkg.rs
│ Qdrant _faces → trace_id │ → tkg_nodes (face_track, etc.)
└──────────────────────────────┘
```
---
## Stage 1: Face Detection
**Script**: `scripts/face_processor.py`
### Flow
1. `swift_face` (Swift/Apple Vision/ANE) → bbox detection per sampled frame
2. `cv2` opens video, crops face from bbox
3. CoreML FaceNet → 512D embedding per face
4. Output: `{uuid}.face.json`
5. Push embeddings to Qdrant `_faces` collection
### Output Format: `{uuid}.face.json`
```json
{
"status": "has_faces",
"frame_count": 563,
"fps": 29.97,
"total_faces": 1200,
"frames": [
{
"frame": 743,
"timestamp": 24.78,
"faces": [
{
"x": 892, // int, pixel
"y": 313, // int, pixel
"width": 78, // int, pixel
"height": 78, // int, pixel
"confidence": 0.733,
"pose_angle": { "angle": "frontal", "roll": 0.77, "yaw": -1.24, "pitch": 0.23 },
"landmarks": { "right_eye": [...], "nose": [...], "left_eye": [...] },
"lips": { "inner_lips": [...], "outer_lips": [...] }
}
]
}
]
}
```
**Key points**:
- bbox is **pixel integer** from Apple Vision, never modified
- face.json uses **list format** (not dict)
- Sampling at ~8Hz (`sample_interval = round(fps / 8)`)
### Qdrant Initial Push
`push_face_embeddings_batch()` in `qdrant_faces.py`:
```python
payload = {
"file_uuid": file_uuid,
"frame": frame_num,
"trace_id": face_idx, # ⚠️ frame-internal index (0, 1, 2...), NOT tracking trace_id
"bbox": {"x": x, "y": y, "width": w, "height": h}, # int pixel
"confidence": 0.5,
"identity_id": None,
"identity_uuid": None,
"stranger_id": None,
}
```
**Important**: `trace_id` at this stage is `face_idx` (index within the frame), used only as a temporary placeholder. It gets overwritten in Stage 2.
---
## Stage 2: Face Tracking
**Scripts**: `scripts/store_traced_faces.py` → `scripts/utils/face_tracker.py`
### Trigger
`job_worker.rs` P2 trigger (line ~1877): after face + asrx processors complete.
```rust
tokio::spawn(async move {
executor.run("store_traced_faces.py", &["--file-uuid", &uuid], ...)
});
```
Skip if `{uuid}.face_traced.json` already exists.
### Flow
1. `store_traced_faces.py` reads `{uuid}.face.json`
2. Converts face.json from list to dict format (frame_num_str → {frame_number, time_seconds, faces})
3. Loads cut boundaries from `{uuid}.cut.json` (if exists)
4. Calls `face_tracker.track_faces(face_data, use_embedding=False, cut_boundaries=...)`
5. Writes `{uuid}.face_traced.json`
6. Calls `update_trace_ids(file_uuid, trace_mapping)` to update Qdrant
### `face_tracker.py:track_faces()`
**Algorithm** (IoU-only, no embedding):
```
For each frame (sorted):
For each face in current frame:
Match against previous frame faces:
- Calculate IoU
- Calculate bbox center distance
- Reject if area ratio > 5x (different zoom level)
- Reject if at-edge → not-at-edge transition (person exited)
If match found → same trace_id as matched face
If no match → new trace_id (next_trace_id++)
Scene cut boundary between frames → force all new traces
```
**Matching conditions** (IoU-only mode):
- IoU > 0.5 AND IoU > 0.35 + distance < 100px → match
- IoU > 0.5 + similarity > 0.65 → match (similarity not used but condition exists)
- similarity > 0.85 → match (not used in IoU-only mode)
- Scene cut boundary → all new traces
### Output Format: `{uuid}.face_traced.json`
Same structure as face.json, but:
- Format converted to **dict** (`frames[str(frame_num)]` → face data)
- Each face gains `trace_id` field (integer)
- Top-level `traces` dict with per-trace statistics
- `metadata.tracking_method = "iou_only"`
- `metadata.traced_at = ISO timestamp`
```json
{
"metadata": {
"fps": 29.97,
"total_frames": 43977,
"tracking_method": "iou_only",
"trace_stats": {
"total_traces": 107,
"active_traces": 107,
"long_traces": 95
}
},
"frames": {
"743": {
"frame_number": 743,
"faces": [
{ "x": 892, "y": 313, "width": 78, "height": 78, "trace_id": 0, ... }
]
}
},
"traces": {
"0": {
"trace_id": 0,
"start_frame": 743,
"end_frame": 783,
"duration_frames": 41,
"total_appearances": 11,
"avg_confidence": 0.72
}
}
}
```
### Qdrant Trace Update
`update_trace_ids()` in `qdrant_faces.py`:
1. Scroll all Qdrant `_faces` points for `file_uuid` (with vector + payload)
2. For each point, build `bbox_key = f"{bbox.x}_{bbox.y}_{bbox.width}_{bbox.height}"`
3. Look up `trace_mapping[frame][bbox_key]` from face_traced.json
4. If match found → set `payload["trace_id"] = real_trace_id`
5. PUT updated points back to Qdrant
**Matching key**: `frame` + `bbox_key` (pixel integer string)
---
## Stage 3: Trace Profile
**Script**: `scripts/backfill_trace_profiles.py`
### Data Source
Qdrant `_faces` collection (source of truth for trace_id assignments).
### Flow
1. Scroll all `_faces` points for each `file_uuid` with `trace_id >= 0`
2. Group by `(file_uuid, trace_id)`
3. For each group:
- `frame_count` = count of points
- `start_frame` = min(frame)
- `end_frame` = max(frame)
- `representative_frame` = frame with max(confidence)
- `representative_bbox` = bbox at representative frame
4. Extract `key_frame.jpg` via ffmpeg at representative frame
5. Crop `key_face.jpg` from key_frame using representative bbox
6. Write `output/{uuid}/trace_{N}/trace_profile.json`
### Output: `output/{uuid}/trace_{N}/trace_profile.json`
```json
{
"version": "1.0",
"file_uuid": "d8acb03870f0cc9b14e01f14a7bf24d6",
"trace_id": 37,
"label": "",
"frame_count": 38,
"start_frame": 1859,
"end_frame": 2100,
"avg_confidence": 0.754,
"key_frame": "key_frame.jpg",
"key_face": "key_face.jpg",
"status": "pending"
}
```
### File Layout
```
output/{uuid}/
trace_0/
trace_profile.json
key_frame.jpg
key_face.jpg
trace_1/
trace_profile.json
key_frame.jpg
key_face.jpg
...
```
---
## Stage 4: TKG Node Construction
**File**: `src/core/processor/tkg.rs`
Reads trace_id from Qdrant `_faces` payload to build knowledge graph nodes:
- `face_track` nodes: one per trace
- `gaze_track`, `lip_track`: linked to face_track via frame alignment
- `co_occurrence` edges: traces that appear in same frame
---
## Qdrant `_faces` Collection Schema
| Field | Type | Description |
|-------|------|-------------|
| `file_uuid` | string | Video file identifier |
| `frame` | int | Video frame number (absolute, not sampled) |
| `trace_id` | int | Face tracking ID (set by Stage 2) |
| `bbox` | `{x, y, width, height}` | Pixel integer coordinates |
| `confidence` | float | Detection confidence |
| `identity_id` | int? | Identity binding (set by identity agent) |
| `identity_uuid` | string? | Identity UUID |
| `stranger_id` | int? | Stranger classification |
**Point ID**: `generate_point_id(file_uuid, frame, face_idx)` — deterministic hash.
---
## Known Issues
### bfba056f5021e2404b0870cc0b1fa851
- **Qdrant**: trace_id = 0,1,2 (face_idx, never updated)
- **face_traced.json**: trace_id = 0-8209 (8210 traces, iou_only)
- **Root cause**: `face_processor.py` re-ran after `store_traced_faces.py`, pushing fresh embeddings with `trace_id=face_idx`, overwriting the updated trace_ids
- **Other 12 files**: all correct
### `update_trace_ids` bbox matching
Matching is by exact `frame` + `bbox_key` string (`x_y_width_height`). Since bbox is pixel integer from the same source, values are identical across face_traced.json and Qdrant. Mismatch only occurs when face_processor.py re-runs and generates different detection results.
---
## File Inventory (2026-07-22)
| file_uuid | traces (Qdrant) | traces (face_traced) | status |
|-----------|-----------------|----------------------|--------|
| 30affad3... | 52 | 53 | ✅ |
| 31a6b821... | 31 | 36 | ⚠️ minor mismatch |
| 352cf73a... | 16 | 25 | ⚠️ minor mismatch |
| 57bd7e43... | 3 | 4 | ✅ |
| 5e5f3de8... | 21 | 22 | ✅ |
| 84d838f2... | 88 | 89 | ✅ |
| 88e72467... | 18 | 19 | ✅ |
| 9cbeb112... | 9 | 17 | ⚠️ minor mismatch |
| bfba056f... | 15 | 8210 | ❌ face_idx not updated |
| c0a9dc37... | 77 | 78 | ✅ |
| c36f3568... | 5601 | 5616 | ⚠️ minor mismatch |
| d8acb038... | 106 | 107 | ✅ |
| fbd82072... | 12 | 13 | ✅ |
---
## Version History
| Version | Date | Changes |
|---------|------|---------|
| 1.0 | 2026-07-22 | Initial document: face detection → tracking → Qdrant → TKG pipeline structure |
+476 -161
View File
@@ -1,198 +1,513 @@
---
document_type: "design_doc"
service: "MOMENTRY_CORE"
title: "File Lifecycle — Pre-Processing & Registration"
version: "V1.2"
date: "2026-05-15"
author: "M5"
status: "draft"
title: File Lifecycle Architecture
version: 1.0
date: 2026-07-22
author: OpenCode
status: Active
scope: File processing pipeline — stages, verification, rebuild
---
# File Lifecycle — Pre-Processing & Registration
# File Lifecycle Architecture V1.0
| Item | Value |
|------|-------|
| Scope | All managed file types (video, image, document, spreadsheet, presentation) |
| Status | Draft |
| Applies to | Pre-process API (explicit) + Register API |
| Key concept | Two-phase flow: birth certificate (`.pre.json`) → civil registration (DB INSERT) |
| Field | Value |
|-------|-------|
| Scope | Complete file processing lifecycle |
| Status | Active |
| Applies to | Pipeline stages, progress tracking, verification, rebuild |
| Related | `FILE_PROFILE_V1.0.md`, `FACE_TRACKING_PIPELINE_V1.0.md` |
> **Applicable to all managed file types**: video, image, document (pdf, docx, pages, key, numbers), spreadsheet, presentation, and any other file registered in the system. The pre-processor registers any file type found by the watcher. ffprobe is used when applicable; files that ffprobe cannot parse receive minimal filesystem metadata as a fallback.
---
## Metaphor
## 1. Overview
Every registered video file passes through a deterministic pipeline of stages.
Each stage must produce a `.json` (or `.jpg`) artifact on disk.
This enables:
- **Verification**: Check pipeline completeness by inspecting artifact existence
- **Rebuild**: Re-run any stage from its input artifacts without re-running the entire pipeline
- **Progress tracking**: Two-layer display (high-level summary + expandable sub-stages)
### Design Principles
1. **Every stage has a `.json` output** — no silent DB-only writes
2. **Any stage can be rebuilt** from its input artifacts
3. **Frontend reads stages from API** — not hardcoded
4. **Verification is disk-first** — check `.json` exists, then validate content, then check DB/Qdrant consistency
5. **Processors are not modified** — this document defines tracking/verification/rebuild only
---
## 2. Stage Architecture
### 2.1 High-Level Stages (6)
| # | Stage | Weight | Sub-Stages | Description |
|---|-------|--------|------------|-------------|
| S0 | Register | 5% | 4 | File metadata + audio track + key frame extraction |
| S1 | Processors | 40% | 8 | Individual processor execution |
| S2 | Post-Process | 20% | 4 | Face trace, Rule1, Vectorize, Identity Agent |
| S3 | TKG Build | 20% | 2 | Temporal Knowledge Graph nodes + edges |
| S4 | Rule2 | 10% | 1 | Relationship chunk ingestion |
| S5 | Complete | 5% | 1 | Final status update |
### 2.2 Sub-Stages (15)
```
SHA256 = DNA or fingerprint (immutable biometric identity)
file mtime = birth moment (preserved by rsync across systems)
birthday (file_uuid anchor) = mtime timestamp
.pre.json = birth certificate
POST /api/v1/files/register = civil registration
status = registered = citizenship completed
S0: Register (5%)
├─ 0a: probe → probe.json
├─ 0b: audio_track → DB: audio_track column (no disk artifact)
├─ 0c: profile → profile.json
└─ 0d: key_frame → key_frame.jpg
S1: Processors (40%)
├─ 1a: cut → cut.json
├─ 1b: asr → asr.json
├─ 1c: asrx → asrx.json (depends: 1a + 1b)
├─ 1d: ocr → ocr.json
├─ 1e: face → face.json (+ Qdrant _faces initial)
├─ 1f: pose → pose.json (depends: 1e)
├─ 1g: appearance → appearance.json (depends: 1f)
└─ 1h: face_dedup → face_cluster.json (depends: 1e) [OPTIONAL + MANUAL]
S2: Post-Process (20%)
├─ 2a: face_trace → face_traced.json (+ Qdrant trace_id update)
├─ 2b: rule1 → rule1.json (ASRX → sentence chunks)
├─ 2c: vectorize → vectorize.json (embeddings → PG + Qdrant)
└─ 2d: identity_agent → identity_agent.json (optional)
S3: TKG Build (20%)
├─ 3a: tkg_nodes → tkg_nodes.json
└─ 3b: tkg_edges → tkg_edges.json
S4: Rule2 (10%)
└─ 4a: rule2 → rule2.json (relationship chunks)
S5: Complete (5%)
└─ 5a: complete → status = "completed"
```
## Two-Phase Flow
A file enters the system in two distinct phases:
| Phase | Action | Analogy | Automatic? | Status |
|-------|--------|---------|:----------:|:------:|
| **Birth** | Pre-process: SHA256 + probe + file_uuid | 出生 + 醫院開出生證明 | ✅ Watcher | `unregistered` |
| **Citizenship** | Register: INSERT into DB | 戶政事務所登記 | ❌ User API | `registered` |
## Phase 1: Pre-Processing (Birth)
### Trigger
Pre-processing is triggered explicitly via the register API or a dedicated pre-process endpoint. It is NOT automatic — the watcher only detects new files without modifying them.
### Computation Steps
### 2.3 Dependency Graph
```
1. fs::metadata(path).modified()
→ birthday = file modification time (mtime, RFC 3339; preserved by rsync -a across systems)
2. SHA256(full file, streaming 64KB chunks)
→ content_hash = 512-bit hex string (file DNA / fingerprint)
3. ffprobe (or minimal fs metadata fallback for non-video)
→ probe_json
4. compute_birth_uuid(mac, birthday, canonical_path, filename)
→ file_uuid = SHA256(mac | birthday | path | filename)[0:32]
5. Write {OUTPUT_DIR}/{file_uuid}.pre.json
S0 (Register)
└─→ S1 (Processors)
├─ 1a (CUT) ─────┐
├─ 1b (ASR) ─────┤
│ └─→ 1c (ASRX) ──→ 2b (Rule1)
├─ 1d (OCR) ──────────────────────→ 3a (TKG Nodes)
├─ 1e (Face) ──┬─→ 1f (Pose) ──→ 1g (Appearance) ──→ 3a
│ ├─→ 1h (FaceDedup) [manual]
│ └─→ 2a (Face Trace) ──→ 3a
└─────────────────────────────────────→ 3a
│
S2: 2c (Vectorize) ←── DB chunks │
S2: 2d (IdentityAgent) ←── face_clusters │
↓
3b (TKG Edges)
│
↓
4a (Rule2)
│
↓
5a (Complete)
```
### Output: `.pre.json` Schema
---
Stored alongside other processor outputs:
## 3. I/O Specification
### 3.1 Register (S0)
| Sub-Stage | Input | Output Artifact | DB Tables | Qdrant |
|-----------|-------|----------------|-----------|--------|
| 0a: probe | video file on disk | `{uuid}.probe.json` | — | — |
| 0b: audio_track | probe.json, video file | DB column only | videos.audio_track | — |
| 0c: profile | probe.json | `{uuid}.profile.json` | videos (INSERT/UPDATE) | — |
| 0d: key_frame | probe.json | `{uuid}.key_frame.jpg` | — | — |
**Audio Track Classification** (S0b):
| Classification | Condition | ASR Behavior |
|----------------|-----------|--------------|
| `no_audio` | No audio track in video | Skip ASR, output `{"status": "no_audio"}` |
| `silent_audio` | Audio track exists but no speech detected | Skip ASR, output `{"status": "silent_audio"}` |
| `music_only` | Audio with no speech (music/sound effects) | Skip ASR, output `{"status": "music_only"}` |
| `speech_only` | Audio with speech only (≥30% speech ratio) | Run ASR normally |
| `speech_with_music` | Speech with background music (<30% speech ratio) | Run ASR normally |
### 3.2 Processors (S1)
| Sub-Stage | Input Artifacts | Output Artifact | DB Tables | Qdrant |
|-----------|----------------|----------------|-----------|--------|
| 1a: cut | probe.json | `{uuid}.cut.json` + `{uuid}_scene_{n}.jpg` | processor_results | — |
| 1b: asr | video file | `{uuid}.asr.json` | processor_results | — |
| 1c: asrx | cut.json, asr.json | `{uuid}.asrx.json` | speaker_detections | — |
| 1d: ocr | video file | `{uuid}.ocr.json` | processor_results | — |
| 1e: face | video file | `{uuid}.face.json` | processor_results | `_faces` (initial push) |
| 1f: pose | face.json, video file | `{uuid}.pose.json` | processor_results | — |
| 1g: appearance | pose.json, video file | `{uuid}.appearance.json` | processor_results | — |
| 1h: face_dedup | face.json | `{uuid}.face_cluster.json` | face_clusters | — |
**Note**: 1h (Face Deduplication) is currently `optional + manual`. It will be integrated into the automated pipeline after testing is complete.
**Scene Key Frames** (1a post-process):
After CUT completes, extracts the middle frame from each scene as `{uuid}_scene_{n}.jpg` for VLM analysis:
| Output | Purpose |
|---------|---------|
| `{uuid}_scene_1.jpg` | Representative frame from scene 1 |
| `{uuid}_scene_2.jpg` | Representative frame from scene 2 |
| ... | ... |
These key frames enable:
- VLM scene understanding (caption, objects, actions)
- Scene-level search and filtering
- Thumbnail generation for scene navigation
### 3.3 Post-Process (S2)
| Sub-Stage | Input Artifacts | Output Artifact | DB Tables | Qdrant |
|-----------|----------------|----------------|-----------|--------|
| 2a: face_trace | face.json | `{uuid}.face_traced.json` | — | `_faces` (trace_id update) |
| 2b: rule1 | asrx.json | `{uuid}.rule1.json` | chunk, pre_chunks | — |
| 2c: vectorize | chunk (DB) | `{uuid}.vectorize.json` | chunk_vectors | main collection |
| 2d: identity_agent | face_cluster.json | `{uuid}.identity_agent.json` | file_identities | — |
### 3.4 TKG Build (S3)
| Sub-Stage | Input Artifacts | Output Artifact | DB Tables | Qdrant |
|-----------|----------------|----------------|-----------|--------|
| 3a: tkg_nodes | All processor JSONs, trace profiles | `{uuid}.tkg_nodes.json` | tkg_nodes | — |
| 3b: tkg_edges | tkg_nodes.json, asrx.json | `{uuid}.tkg_edges.json` | tkg_edges | — |
### 3.5 Rule2 (S4)
| Sub-Stage | Input Artifacts | Output Artifact | DB Tables | Qdrant |
|-----------|----------------|----------------|-----------|--------|
| 4a: rule2 | tkg_edges.json, chunk (DB) | `{uuid}.rule2.json` | chunk (relationship type) | main collection |
### 3.6 Complete (S5)
| Sub-Stage | Input | Output | DB Tables |
|-----------|-------|--------|-----------|
| 5a: complete | All above stages verified | status = "completed" | videos.status |
---
## 4. Verification
### 4.1 Verification Levels
Each sub-stage has three verification levels:
| Level | Check | Description |
|-------|-------|-------------|
| L1: Artifact exists | `{uuid}.{stage}.json` on disk | Required for all stages |
| L2: Content valid | JSON parseable + non-empty array/object | Ensures output is usable |
| L3: DB/Qdrant consistent | Row count > 0 or point count > 0 | Ensures data was written |
### 4.2 Verification Matrix
| Sub-Stage | L1 (exists) | L2 (valid) | L3 (DB/Qdrant) |
|-----------|:-----------:|:----------:|:--------------:|
| 0a: probe | `.probe.json` | non-empty | — |
| 0b: profile | `.profile.json` | has file_uuid | videos row exists |
| 0c: key_frame | `.key_frame.jpg` | file size > 0 | — |
| 1a: cut | `.cut.json` | non-empty | processor_results > 0 |
| 1b: asr | `.asr.json` | non-empty | processor_results > 0 |
| 1c: asrx | `.asrx.json` | non-empty | speaker_detections > 0 |
| 1d: ocr | `.ocr.json` | non-empty | processor_results > 0 |
| 1e: face | `.face.json` | non-empty | Qdrant `_faces` > 0 |
| 1f: pose | `.pose.json` | non-empty | processor_results > 0 |
| 1g: appearance | `.appearance.json` | non-empty | processor_results > 0 |
| 1h: face_dedup | `.face_cluster.json` | non-empty | face_clusters > 0 |
| 2a: face_trace | `.face_traced.json` | non-empty | Qdrant `_faces` trace_id set |
| 2b: rule1 | `.rule1.json` | non-empty | chunk (sentence) > 0 |
| 2c: vectorize | `.vectorize.json` | non-empty | chunk_vectors > 0 |
| 2d: identity_agent | `.identity_agent.json` | non-empty | file_identities > 0 |
| 3a: tkg_nodes | `.tkg_nodes.json` | non-empty | tkg_nodes > 0 |
| 3b: tkg_edges | `.tkg_edges.json` | non-empty | tkg_edges > 0 |
| 4a: rule2 | `.rule2.json` | non-empty | chunk (relationship) > 0 |
### 4.3 Status Values
| Status | Meaning |
|--------|---------|
| `pending` | Not yet started |
| `running` | Currently executing |
| `completed` | L1 + L2 + L3 all pass |
| `failed` | L1 passes but L2 or L3 fails |
| `missing` | L1 fails (artifact not on disk) |
| `skipped` | Optional stage not run |
---
## 5. Rebuild
### 5.1 Rebuild Principle
Any sub-stage can be rebuilt independently:
1. Read input artifacts (from disk or DB)
2. Re-run the stage logic (processor or post-processor)
3. Write output artifact + update DB/Qdrant
### 5.2 Rebuild Dependency
To rebuild stage N, all its dependency stages must be `completed`:
| Stage | Required Dependencies |
|-------|----------------------|
| 0a-0c | video file on disk |
| 1a: cut | 0a (probe) |
| 1b: asr | video file |
| 1c: asrx | 1a (cut) + 1b (asr) |
| 1d: ocr | video file |
| 1e: face | video file |
| 1f: pose | 1e (face) |
| 1g: appearance | 1f (pose) |
| 1h: face_dedup | 1e (face) |
| 2a: face_trace | 1e (face) |
| 2b: rule1 | 1c (asrx) |
| 2c: vectorize | 2b (rule1) — chunks in DB |
| 2d: identity_agent | 1h (face_dedup) — optional |
| 3a: tkg_nodes | 1e (face), 2a (face_trace), 1c (asrx), 1d (ocr), 1g (appearance) |
| 3b: tkg_edges | 3a (tkg_nodes) + 1c (asrx) |
| 4a: rule2 | 3b (tkg_edges) + 2b (rule1) — chunks in DB |
| 5a: complete | All required stages completed |
### 5.3 Rebuild API
```
{OUTPUT_DIR}/
{file_uuid}.probe.json ← ffprobe
{file_uuid}.face.json ← face detection
{file_uuid}.pre.json ← pre-processor (NEW)
POST /api/v1/file/:file_uuid/rebuild/:stage
```
- Validates dependencies are met
- Re-runs the stage
- Returns updated verification status
### 5.4 Rebuild via CLI
```bash
# Check all stages
python3 scripts/lifecycle_check.py --file-uuid <UUID>
# Rebuild specific stage
python3 scripts/lifecycle_check.py --file-uuid <UUID> --rebuild 1c
# Rebuild from first missing stage
python3 scripts/lifecycle_check.py --file-uuid <UUID> --rebuild auto
```
---
## 6. Frontend Display
### 6.1 Two-Layer Architecture
**Layer 1: High-Level Summary** (default view)
```
┌─────────────────────────────────────────────────────────┐
│ ▶ S0: Register ████████████ 3/3 completed │
│ ▶ S1: Processors ████████░░░░ 6/8 partial │
│ ▶ S2: Post-Process ██░░░░░░░░░░ 1/4 running │
│ ▶ S3: TKG Build ░░░░░░░░░░░░ 0/2 pending │
│ ▶ S4: Rule2 ░░░░░░░░░░░░ 0/1 pending │
│ ▶ S5: Complete ░░░░░░░░░░░░ 0/1 pending │
└─────────────────────────────────────────────────────────┘
```
**Layer 2: Expandable Sub-Stages** (click to expand)
```
┌─────────────────────────────────────────────────────────┐
│ ▼ S1: Processors ████████░░░░ 6/8 partial │
│ ├─ 1a: CUT ✅ completed │
│ ├─ 1b: ASR ✅ completed │
│ ├─ 1c: ASRX ✅ completed │
│ ├─ 1d: OCR ✅ completed │
│ ├─ 1e: Face ✅ completed │
│ ├─ 1f: Pose ✅ completed │
│ ├─ 1g: Appearance ❌ missing │
│ └─ 1h: Face Dedup ⏭ skipped (manual) │
└─────────────────────────────────────────────────────────┘
```
### 6.2 Sub-Stage Display Names
| Code Name | Display Name |
|-----------|-------------|
| probe | Probe (ffprobe) |
| profile | File Profile |
| key_frame | Key Frame |
| cut | Scene Detection (CUT) |
| asr | Speech Recognition (ASR) |
| asrx | Speaker Diarization (ASRX) |
| ocr | Text Recognition (OCR) |
| face | Face Detection |
| pose | Pose Estimation |
| appearance | Appearance Features |
| face_dedup | Face Deduplication |
| face_trace | Face Tracking |
| rule1 | Rule1 Ingestion |
| vectorize | Vector Embedding |
| identity_agent | Identity Agent |
| tkg_nodes | TKG Nodes |
| tkg_edges | TKG Edges |
| rule2 | Rule2 Ingestion |
| complete | Complete |
### 6.3 API Contract
The frontend fetches stage data from:
```
GET /api/v1/stats/pipeline/:file_uuid
```
Response:
```json
{
"file_name": "charade.mp4",
"file_path": "/data/demo/charade.mp4",
"canonical_path": "/private/data/demo/charade.mp4",
"content_hash": "a1b2c3d4e5f6...",
"probe_json": {
"format": { "duration": "6879.3", "size": "2147483648" },
"streams": [...]
},
"birthday": "2026-05-15T02:15:00Z",
"file_uuid": "aeed71342a899fe4b4c57b7d41bcb692",
"file_size": 2147483648,
"file_type": "video | image | document | audio",
"pre_processed_at": "2026-05-15T02:15:05Z"
"file_uuid": "abc123",
"overall_progress": 0.45,
"stages": [
{
"name": "register",
"weight": 0.05,
"progress": 1.0,
"status": "completed",
"detail": "3/3 sub-stages",
"sub_stages": [
{"name": "probe", "status": "completed", "artifact": "probe.json"},
{"name": "profile", "status": "completed", "artifact": "profile.json"},
{"name": "key_frame", "status": "completed", "artifact": "key_frame.jpg"}
]
},
{
"name": "processors",
"weight": 0.40,
"progress": 0.75,
"status": "partial",
"detail": "6/8 sub-stages",
"sub_stages": [
{"name": "cut", "status": "completed", "artifact": "cut.json"},
{"name": "asr", "status": "completed", "artifact": "asr.json"},
{"name": "asrx", "status": "completed", "artifact": "asrx.json"},
{"name": "ocr", "status": "completed", "artifact": "ocr.json"},
{"name": "face", "status": "completed", "artifact": "face.json"},
{"name": "pose", "status": "completed", "artifact": "pose.json"},
{"name": "appearance", "status": "missing", "artifact": "appearance.json"},
{"name": "face_dedup", "status": "skipped", "artifact": "face_cluster.json"}
]
}
],
"updated_at": "2026-07-22T18:00:00Z"
}
```
### Key Design: file_uuid = f(mac, birthday, path, filename)
---
The `birthday` is `file modification time` (mtime) — obtained from `fs::metadata().modified()`. Using mtime instead of birthtime ensures file_uuid stability when files are transferred between systems via rsync (which preserves mtime but not birthtime on macOS).
## 7. Weight Distribution
### 7.1 High-Level Stage Weights
| Stage | Weight | Rationale |
|-------|--------|-----------|
| S0: Register | 5% | Fast, prerequisite for everything |
| S1: Processors | 40% | Most time-consuming, GPU-bound |
| S2: Post-Process | 20% | Face trace + Rule1 + Vectorize |
| S3: TKG Build | 20% | Node + edge construction |
| S4: Rule2 | 10% | Relationship chunk creation |
| S5: Complete | 5% | Final status update |
### 7.2 Processor Sub-Weights (within S1 = 40%)
| Processor | Sub-Weight | Rationale |
|-----------|-----------|-----------|
| CUT | 5% | Scene detection, ~10s |
| ASR | 15% | whisper-small, ~2min/10min video |
| ASRX | 20% | Speaker diarization, ~3min |
| OCR | 10% | PaddleOCR, ~1min |
| Face | 15% | CoreML FaceNet, ~1min |
| Pose | 10% | mediapipe, ~1min |
| Appearance | 5% | Feature extraction, ~30s |
| Face Dedup | 0% | Manual (not in automated pipeline) |
---
## 8. Artifact Naming Convention
All artifacts live in the output directory (`MOMENTRY_OUTPUT_DIR`):
```
birthday = 2026-05-15T02:15:00Z ← file birth time, never changes
↓
file_uuid = SHA256(mac | birthday | path | filename)
↓
Same file: same path + filename → same file_uuid, regardless of registration count
Different files: different content_hash → different file_uuid (even if same name)
{output_dir}/
├─ {uuid}.probe.json # S0: ffprobe metadata
├─ {uuid}.profile.json # S0: FileProfile
├─ {uuid}.key_frame.jpg # S0: extracted key frame
├─ {uuid}.cut.json # S1: scene boundaries
├─ {uuid}.asr.json # S1: speech transcription
├─ {uuid}.asrx.json # S1: speaker diarization
├─ {uuid}.ocr.json # S1: text detections
├─ {uuid}.face.json # S1: face detections + embeddings
├─ {uuid}.face_cluster.json # S1: face clustering (optional)
├─ {uuid}.pose.json # S1: pose estimations
├─ {uuid}.appearance.json # S1: appearance features
├─ {uuid}.face_traced.json # S2: face tracking with trace_id
├─ {uuid}.rule1.json # S2: sentence chunks
├─ {uuid}.vectorize.json # S2: embedding stats
├─ {uuid}.identity_agent.json # S2: identity matching (optional)
├─ {uuid}.tkg_nodes.json # S3: TKG node dump
├─ {uuid}.tkg_edges.json # S3: TKG edge dump
├─ {uuid}.rule2.json # S4: relationship chunks
└─ {uuid}/ # Trace profiles directory
├─ trace_0/
│ ├─ trace_profile.json
│ ├─ key_frame.jpg
│ └─ key_face.jpg
├─ trace_1/
│ └─ ...
└─ trace_N/
```
## Phase 2: Registration (Citizenship)
---
### POST /api/v1/files/register
## 9. Current State Audit (Gamma 8)
```bash
curl -X POST http://localhost:3002/api/v1/files/register \
-H "X-API-Key: ..." \
-H "Content-Type: application/json" \
-d '{"file_path":"/data/demo/charade.mp4"}'
```
File: `d3f9ae8e471a1fc4d47022c66091b920` (Gamma 8-Director Chih-Lin Yang)
### Flow
| Sub-Stage | Artifact | Status |
|-----------|----------|--------|
| 0a: probe | probe.json | ✅ exists |
| 0b: profile | profile.json | ❌ missing |
| 0c: key_frame | key_frame.jpg | ❌ missing |
| 1a: cut | cut.json | ✅ exists |
| 1b: asr | asr.json | ✅ exists |
| 1c: asrx | asrx.json | ✅ exists |
| 1d: ocr | ocr.json | ✅ exists |
| 1e: face | face.json | ✅ exists |
| 1f: pose | pose.json | ✅ exists |
| 1g: appearance | appearance.json | ❌ missing |
| 1h: face_dedup | face_cluster.json | ⏭ skipped (manual) |
| 2a: face_trace | face_traced.json | ✅ exists |
| 2b: rule1 | rule1.json | ❌ missing |
| 2c: vectorize | vectorize.json | ❌ missing |
| 2d: identity_agent | identity_agent.json | ❌ missing |
| 3a: tkg_nodes | tkg_nodes.json | ❌ missing |
| 3b: tkg_edges | tkg_edges.json | ❌ missing |
| 4a: rule2 | rule2.json | ❌ missing |
```
1. Check {OUTPUT_DIR}/{file_uuid}.pre.json
├─ Exists AND content_hash matches → use cached (skip SHA256 + probe)
└─ Not exists OR hash mismatch → compute fresh (existing logic)
2. Dedup check: SELECT file_uuid FROM videos WHERE content_hash = $1
├─ Found → already_exists: true (identical DNA = same file)
└─ Not found → continue
3. Name conflict check + auto-rename if needed
└─ charade.mp4 → charade (1).mp4 (same name, different content)
4. INSERT INTO videos (
file_uuid, file_path, file_name, file_type,
duration, width, height, fps,
probe_json, content_hash, status, registration_time
) VALUES (
$1, $2, $3, $4, $5, $6, $7, $8, $9, $10,
'registered', NOW() ← status=registered, registration_time=NOW()
)
```
## Data Separation
| Field | Source | Computed When | Mutable |
|-------|--------|---------------|:------:|
| `birthday` | `fs::metadata().modified()` (mtime) | Pre-process (once) | ❌ Never (stable across rsync) |
| `content_hash` (SHA256) | Full file | Pre-process (once) | ❌ Never (unless file modified) |
| `file_uuid` | SHA256(mac\|birthday\|path\|filename) | Pre-process (once) | ❌ Never |
| `registration_time` | `NOW()` at register | Register API | ✅ Per registration |
| `status` | — | Register API | `unregistered` → `registered` |
## File Lifecycle State Diagram
```
File detected by watcher (detection only, no modification)
│
│ Pre-processing triggered explicitly (API or register)
▼
[Pre-Processor]
├─ SHA256 (DNA / fingerprint)
├─ ffprobe (metadata extraction)
└─ file_uuid (birth certificate ID)
│
▼
{file_uuid}.pre.json
status = unregistered (no DB record)
│
│ (user calls POST /api/v1/files/register)
▼
[Register Handler]
├─ Read .pre.json → skip recomputation
├─ Dedup check (content_hash collision?)
├─ Name check + rename?
└─ INSERT INTO videos
│
▼
status = registered
registration_time = NOW()
```
## Implementation Checklist
| # | Task | File |
|---|------|------|
| 1 | Expose `pre_process_file()` as public function (SHA256 + probe + file_uuid → `.pre.json`) | `src/watcher/watcher.rs` |
| 2 | Register: read `.pre.json`, skip SHA256/probe if cached | `src/api/server.rs` → `register_single_file` |
| 3 | file_uuid: use `birthday` from `.pre.json` (or `fs::metadata().modified()` fallback) | `src/api/server.rs` |
| 4 | INSERT status: `registered`, registration_time: `NOW()` | `src/api/server.rs` |
**Observations**:
- S1 processors mostly complete, but Appearance missing (1g)
- S0 profile/key_frame missing (registration may not have created them)
- S2-S4 all have DB data but no flat `.json` dumps
---
## Version History
| Version | Date | Changes |
|---------|------|---------|
| V1.0 | 2026-05-15 | Initial design — birth certificate (pre-process) + civil registration two-phase flow |
| V1.1 | 2026-05-15 | Reclassified from DESIGN to STANDARDS as design standard |
| V1.2 | 2026-05-15 | mtime replaces birthtime for file_uuid stability across rsync; watcher is detection-only |
| Version | Date | Author | Changes |
|---------|------|--------|---------|
| 1.2 | 2026-07-22 | OpenCode | Added CUT scene key frames extraction for VLM analysis |
| 1.1 | 2026-07-22 | OpenCode | Added S0b: audio_track classification (VAD) — 6 stages, 15 sub-stages |
| 1.0 | 2026-07-22 | OpenCode | Initial design — 6 stages, 14 sub-stages, I/O specs, verification, rebuild |
+153
View File
@@ -0,0 +1,153 @@
# File Profile V1.0
**Status:** Active
**Version:** 1.0
**Date:** 2026-07-22
**Scope:** File identity artifact — persistent on-disk profile per registered file
---
## Problem
- 5 zombie files in DB: `file_uuid` exists but `file_name` and `file_path` are empty — no way to recover
- File identity lives only in PostgreSQL; no on-disk fallback
- `birth_registration` written by `ingestion.rs` but **not** by `files.rs` API path
- No file history — if a file moves or is renamed, no record of where it was
## Design
A JSON file created **first** during registration, stored flat in `MOMENTRY_OUTPUT_DIR`:
```
{MOMENTRY_OUTPUT_DIR}/{file_uuid}.profile.json
```
### JSON Structure
```json
{
"version": "1.0",
"file_uuid": "84d838f260e1881a0daa55fabbc8e434",
"file_name": "view28.mp4",
"file_type": "video",
"birth": {
"mac_address": "a1:b2:c3:d4:e5:f6",
"birthday": "2026-04-13T23:00:49+08:00",
"original_path": "/Users/accusys/momentry/var/sftpgo/data/demo",
"original_filename": "view28.mp4",
"canonical_path": "/Users/accusys/momentry/var/sftpgo/data/demo/view28.mp4",
"content_hash": "abc123..."
},
"current": {
"path": "/Users/accusys/momentry/var/sftpgo/data/demo/view28.mp4",
"file_name": "view28.mp4",
"file_type": "video"
},
"history": [
{
"action": "registered",
"timestamp": "2026-07-22T14:30:00+08:00",
"path": "/Users/accusys/momentry/var/sftpgo/data/demo/view28.mp4",
"file_name": "view28.mp4"
}
],
"metadata": {
"duration": 243.24,
"width": 720,
"height": 890,
"fps": 60.0,
"total_frames": 7297
},
"key_frame": null
}
```
### Fields
| Field | Purpose |
|-------|---------|
| `version` | Profile schema version (for future migration) |
| `file_uuid` | The deterministic UUID |
| `file_name` | Original filename at registration |
| `file_type` | video/audio/image/document/... |
| `birth.mac_address` | MAC address used to compute UUID |
| `birth.birthday` | File mtime (RFC3339) used to compute UUID |
| `birth.original_path` | Parent directory at registration |
| `birth.original_filename` | Filename at registration |
| `birth.canonical_path` | Canonical (resolved symlinks) path at registration |
| `birth.content_hash` | SHA256 of file content |
| `current.path` | Latest known path (updated when file moves) |
| `current.file_name` | Latest known filename (updated on rename) |
| `current.file_type` | Latest file type |
| `history` | Array of all path/name changes with timestamps |
| `metadata` | Media info (duration, resolution, etc.) |
| `key_frame` | Base64-encoded JPEG of representative frame (video only), or null |
### key_frame
For video files, `key_frame` stores a **base64-encoded JPEG** of a representative frame extracted at registration time (typically at 10% of duration or the first non-black frame). For non-video files, this field is `null`.
Purpose:
- Instant visual identification without needing to open the video
- Fallback if thumbnails or `.faces/` crops are deleted
- Portable — the profile file is self-contained
Extraction:
- Uses ffmpeg to grab a frame at `duration * 0.1` (or first frame if duration unknown)
- JPEG quality 85, max width 640px
- Stored inline as base64 string in the JSON
## Implementation
### New Module
`src/core/file_profile.rs` — `FileProfile` struct with:
- `from_registration_params(...)` — build at registration time
- `load_from_disk(uuid, output_dir)` — read `{uuid}.profile.json`
- `save_to_disk(&self, output_dir)` — write `{uuid}.profile.json`
- `update_current_path(&mut self, new_path, new_name)` — append to history, update current
- `extract_key_frame(video_path, duration)` — ffmpeg frame extraction + base64
### Registration Flow
1. DB INSERT (existing)
2. **Build FileProfile** from all params (mac, birthday, path, name, content_hash, probe metadata)
3. **Extract key_frame** if video (ffmpeg)
4. **Save `{uuid}.profile.json`** — first artifact on disk
5. CUT processing (existing)
### Update Flow
When `file_path` or `file_name` changes via API:
1. Load profile from disk
2. `profile.update_current_path(new_path, new_name)`
3. Save updated profile
### Fallback Flow
If DB data is missing (zombie files):
1. Load profile from disk
2. Profile's `current.path` and `current.file_name` provide recovery data
### Files Changed
| File | Change |
|------|--------|
| `src/core/file_profile.rs` | **NEW** — FileProfile struct |
| `src/core/mod.rs` | Add `pub mod file_profile` |
| `src/api/files.rs` | Write profile after registration; fallback; enrich GET; cleanup |
| `src/core/ingestion.rs` | Write profile after registration |
| `src/api/profile.rs` | Enrich GET file-profile with profile data + history |
### Backfill
Existing 16 files get profiles generated from DB fields + probe.json data.
5 zombie files get minimal profiles (UUID + file_type + content_hash from DB).
---
## Version History
| Version | Date | Change |
|---------|------|--------|
| 1.0 | 2026-07-22 | Initial design — file profile with key_frame |
@@ -0,0 +1,339 @@
# Face-Pose-Appearance Tracking Design
**Version**: 1.0
**Date**: 2026-07-19
**Status**: Ready for Implementation
---
## 1. Overview
本文件定義 Face、Pose、Appearance 的追蹤系統設計,包含:
- Trace ID 繼承規則
- 擴張邏輯
- Appearance 色塊提取
- Agent Search 整合
---
## 2. Core Concepts
### 2.1 Identity vs Tracking
| Processor | Purpose | Description |
|-----------|---------|-------------|
| **Face** | Identity | Who is this person? 需要高品質 embedding |
| **Pose** | Tracking | Where is this person? 當 face occluded 時維持追蹤 |
| **Appearance** | Tracking | What do they look like? 當 pose occluded 時維持追蹤 |
### 2.2 Offline Processing Advantage
Offline 處理可以先做 face detection,再從 face traces 擴張 pose/appearance:
```
Face Detection → 知道身份錨點
↓
Face Tracking → 給予 trace_id
↓
Pose Expansion → 從 face traces 向外擴張
↓
Appearance Expansion → 從 pose traces 向外擴張
```
---
## 3. Processing Pipeline
### 3.1 Pipeline Order
| Order | Processor | Dependencies | Output | Description |
|-------|-----------|-------------|--------|-------------|
| 1 | `cut` | — | cut.json | Scene detection |
| 2 | `face` | — | face.json | Face detection (8Hz) + embedding |
| 3 | `face_trace` | face | face_traced.json | Face tracking (IoU + embedding) |
| 4 | `pose` | face_trace | pose.json | Pose expansion from traces |
| 5 | `appearance` | pose | appearance.json | Appearance extraction |
| 6 | `asr` | cut | asr.json | Speech-to-text |
| 7 | `asrx` | asr | asrx.json | Speaker diarization |
### 3.2 Sampling Rate
- **公式**: `sample_interval = floor(fps / 8)`
- **確保**: ≥ 8Hz 取樣率
- **範例**:
- 24fps → interval = 3 → 8Hz
- 30fps → interval = 3 → 10Hz
- 60fps → interval = 7 → 8.6Hz
---
## 4. Trace ID Inheritance
### 4.1 Inheritance Chain
```
Face Trace (identity anchor)
│ trace_id = 1, 2, 3, ...
│
▼ inherits trace_id
Pose Expansion
│ same trace_id per person
│
▼ inherits trace_id
Appearance Expansion
│ same trace_id per person
```
### 4.2 Frame Count Relationship
```
face_frames ≤ pose_frames ≤ appearance_frames
```
**原因**:
- Face: 只有臉部可見的 frames
- Pose: Face frames + 擴張 frames(臉被遮但身體可見)
- Appearance: Pose frames + 擴張 frames
### 4.3 Trace Connection
```
Face trace A (frames 1-10) Face trace B (frames 20-30)
↘ ↙
Pose 連接 (frames 15-18)
(同一人,中間臉被遮住)
```
**意義**: Pose 可以連接斷開的 face traces,屬於同一人。
---
## 5. Expansion Rules
### 5.1 Pose Expansion
**Algorithm**:
1. 讀取 face_traced.json,取得每個 trace_id 的 frames
2. 對每個 trace 的 frames 向外擴張(逐幀檢查)
3. 連續 3 幀無 pose detection → 停止擴張
4. 繼承 trace_id
5. 輸出 8Hz 取樣
**Parameters**:
| Parameter | Value | Description |
|-----------|-------|-------------|
| `miss_threshold` | 3 | 連續無檢測幀數 |
| `output_rate` | 8Hz | 輸出取樣率 |
### 5.2 Appearance Expansion
**Algorithm**:
1. 讀取 pose.json,取得每個 trace_id 的 frames
2. 對每個 pose frame,在 keypoint 位置提取顏色
3. 記錄整體亮度
4. 輸出 8Hz 取樣
**Parameters**:
| Parameter | Value | Description |
|-----------|-------|-------------|
| `color_radius` | 15 | 顏色取樣半徑(pixels) |
| `output_rate` | 8Hz | 輸出取樣率 |
---
## 6. Pose Output
### 6.1 Data Structure
```json
{
"frame_count": 1000,
"fps": 24.0,
"frames": [
{
"frame": 100,
"timestamp": 4.16,
"trace_id": 1,
"persons": [
{
"keypoints": [
{"name": "nose", "x": 315.9, "y": 364.2, "confidence": 0.85},
{"name": "left_shoulder", "x": 290.0, "y": 400.0, "confidence": 0.92}
],
"bbox": {"x": 280, "y": 350, "width": 100, "height": 200}
}
]
}
]
}
```
### 6.2 Bbox Validation
**原則**: Face bbox 應在 Pose bbox 內,或 IoU > 0.5
```
┌─────────────────────────┐
│ Pose bbox │
│ ┌─────────┐ │
│ │ Face │ │
│ │ bbox │ │
│ └─────────┘ │
└─────────────────────────┘
```
**用途**:
- 品質驗證:確保 pose/face 屬於同一人
- 匹配追蹤:用 bbox overlap 匹配 face/pose
---
## 7. Appearance Output
### 7.1 Keypoint-based Color Extraction
**原理**: 在 pose keypoint 位置取周圍平均色
```
Pose Keypoints 座標
↓
在每個 keypoint 位置取色
↓
記錄為 appearance
```
### 7.2 Body Part Mapping
| Keypoints | Body Part | Description |
|-----------|-----------|-------------|
| nose, eyes, ears | head | 帽子、頭髮顏色 |
| shoulders | torso | 上衣顏色 |
| hips, knees | legs | 褲子顏色 |
| ankles | feet | 鞋子顏色 |
### 7.3 Data Structure
```json
{
"frame_count": 1000,
"fps": 24.0,
"frames": [
{
"frame": 100,
"timestamp": 4.16,
"trace_id": 1,
"brightness": 0.75,
"colors": {
"head": [180, 150, 120],
"torso": [255, 50, 50],
"legs": [50, 50, 200],
"feet": [50, 200, 50]
}
}
]
}
```
### 7.4 Lighting Record
```json
{
"brightness": 0.75
}
```
**用途**: 不同光源下的顏色校正
---
## 8. VLM Complementary Strategy
### 8.1 Two-Level Approach
| Level | Method | Purpose |
|-------|--------|---------|
| **L1** | Keypoint 快取色 | 快速搜尋、初步候選 |
| **L2** | VLM 驗證 | 複雜情況、細節補充(可選) |
### 8.2 Workflow
```
搜尋「穿紅上衣的人」
↓
L1: Keypoint 取色搜尋 → Top 20 候選
↓
L2: VLM 驗證(需要時)→ 確認顏色、補充細節
↓
最終結果 → Top 10 + 置信度
```
---
## 9. Agent Search Integration
### 9.1 Agent Tool Design
```python
def search_by_appearance(
color: str, # "red", "blue", "green"
body_part: str, # "torso", "legs", "feet"
top_k: int = 10
) -> List[SearchResult]:
"""
搜尋穿特定顏色衣物的人
Returns:
[
{"trace_id": 1, "identity": "John", "confidence": 0.85},
{"trace_id": 2, "identity": "Mary", "confidence": 0.72},
]
"""
```
### 9.2 Query Examples
| User Query | Agent Action |
|------------|--------------|
| 「穿紅上衣的人是誰?」 | search_by_appearance("red", "torso") → match identity |
| 「穿綠鞋子的人」 | search_by_appearance("green", "feet") |
| 「戴黑帽子的人」 | search_by_appearance("black", "head") |
### 9.3 Top-K Strategy
- **原則**: 找 top 10-20 最相似的
- **容許誤差**: 光源、角度差異可接受
- **近似即可**: 不需精確匹配
---
## 10. Implementation Files
| Component | File | Status |
|-----------|------|--------|
| Face Detection | `swift_face.swift` | ✅ Complete |
| Face Tracking | `store_traced_faces.py` | ✅ Complete |
| Pose Expansion | `swift_pose_expansion.swift` | ✅ Complete |
| Appearance Expansion | `swift_appearance_expansion.swift` | ✅ Complete |
| Pose Processor | `pose_processor_v2.py` | ✅ Complete |
| Appearance Processor | `appearance_processor_v2.py` | ✅ Complete |
---
## 11. Testing Checklist
- [ ] 清除測試檔案重新註冊
- [ ] 執行完整流程:face → trace → pose → appearance
- [ ] 驗證 trace_id 繼承正確性
- [ ] 驗證 frame count 關係 (face ≤ pose ≤ appearance)
- [ ] 驗證 bbox 包含關係 (face bbox ⊂ pose bbox)
- [ ] 測試 Agent search_by_appearance
---
## 12. Version History
| Version | Date | Changes |
|---------|------|---------|
| 1.0 | 2026-07-19 | Initial design |
+209
View File
@@ -0,0 +1,209 @@
# Trace ID Inheritance & Expansion Rules
**Date**: 2026-07-19
**Author**: Core Team
**Status**: Final
---
## Overview
This document defines the trace ID inheritance rules and expansion logic for Face, Pose, and Appearance processing.
---
## Core Concepts
### Face = Identity, Pose/Appearance = Tracking
| Processor | Purpose | Description |
|-----------|---------|-------------|
| **Face** | Identity | Who is this person? Requires high-quality embedding for recognition. |
| **Pose** | Tracking | Where is this person? Maintains tracking when face is occluded. |
| **Appearance** | Tracking | What do they look like? Maintains tracking when pose is occluded. |
### Offline Processing Advantage
In offline processing, we can:
1. First detect all faces (identity anchors)
2. Then expand pose/appearance from face traces
This is different from real-time tracking where pose/appearance runs continuously and face anchors identity when visible.
---
## Processing Pipeline
### Step 1: Face Detection (8Hz)
```
swift_face → face.json
```
- Sampling rate: `floor(fps / 8)` (ensures ≥ 8Hz)
- Output: Face bounding boxes with landmarks and embeddings
### Step 2: Face Tracking
```
store_traced_faces.py → face_traced.json
```
- Algorithm: IoU + embedding similarity
- Output: Each face assigned a `trace_id`
- Purpose: Group same-person faces across frames
### Step 3: Pose Expansion
```
swift_pose_expansion → pose.json
```
**Input**: face_traced.json (frames with trace_id)
**Expansion Algorithm**:
1. For each trace_id, get all face frames
2. Expand outward (forward/backward) checking for pose
3. Stop when 3 consecutive frames have no pose detection
4. Inherit trace_id from face
**Output**: Pose keypoints with inherited trace_id
### Step 4: Appearance Expansion
```
swift_appearance_expansion → appearance.json
```
**Input**: pose.json (frames with trace_id)
**Expansion Algorithm**:
1. For each trace_id, get all pose frames
2. Expand outward (forward/backward) checking for appearance
3. Stop when 3 consecutive frames have HSV similarity < 0.5
4. Inherit trace_id from pose
**Output**: HSV histograms with inherited trace_id
---
## Trace ID Inheritance
```
Face Trace (identity anchor)
│
│ inherits trace_id
▼
Pose Expansion
│
│ inherits trace_id
▼
Appearance Expansion
```
**Key Points:**
- Trace ID originates from face tracking
- Pose inherits the same trace_id (same person)
- Appearance inherits the same trace_id (same person)
- This enables linking all detections to the same identity
---
## Frame Count Relationship
```
face frames ≤ pose frames ≤ appearance frames
```
**Explanation:**
- Face: Only frames where face is clearly visible
- Pose: Face frames + expanded frames (pose may still be visible when face is occluded)
- Appearance: Pose frames + expanded frames (appearance may still be visible when pose is occluded)
---
## Expansion Rules
### Pose Expansion
| Parameter | Value | Description |
|-----------|-------|-------------|
| `miss_threshold` | 3 | Stop after 3 consecutive frames without pose |
| `max_range` | 300 frames | Maximum expansion distance (≈10s at 30fps) |
| `output_rate` | 8Hz | Output sampling rate |
### Appearance Expansion
| Parameter | Value | Description |
|-----------|-------|-------------|
| `miss_threshold` | 3 | Stop after 3 consecutive frames with similarity < 0.5 |
| `similarity_threshold` | 0.5 | HSV histogram similarity threshold |
| `max_range` | 300 frames | Maximum expansion distance |
| `output_rate` | 8Hz | Output sampling rate |
---
## Tracking Continuity
### Pose Can Connect Face Traces
```
Face trace A (frames 1-10) Face trace B (frames 20-30)
↘ ↙
Pose connects (frames 15-18)
(Same person, face was occluded)
```
When pose expansion from two face traces overlaps, they may belong to the same person. Future enhancement: pose-based trace merging.
### Appearance Can Connect Pose Traces
Similar to pose, appearance similarity can connect pose traces when pose is temporarily occluded.
---
## Implementation Files
| Component | File |
|-----------|------|
| Face Detection | `scripts/swift_processors/swift_face.swift` |
| Face Tracking | `scripts/store_traced_faces.py` |
| Pose Expansion | `scripts/swift_processors/swift_pose_expansion.swift` |
| Appearance Expansion | `scripts/swift_processors/swift_appearance_expansion.swift` |
| Pose Processor Wrapper | `scripts/pose_processor_v2.py` |
| Appearance Processor Wrapper | `scripts/appearance_processor_v2.py` |
| Dependencies Definition | `src/core/db/postgres_db.rs:568-577` |
---
## Testing
### Verify Trace ID Inheritance
```bash
# Check face traces
cat /Users/accusys/momentry/output/$FILE_UUID.face_traced.json | jq '.frames[].faces[].trace_id' | sort | uniq -c
# Check pose traces (should have same trace_ids)
cat /Users/accusys/momentry/output/$FILE_UUID.pose.json | jq '.frames[].trace_id' | sort | uniq -c
# Check appearance traces (should have same trace_ids)
cat /Users/accusys/momentry/output/$FILE_UUID.appearance.json | jq '.frames[].trace_id' | sort | uniq -c
```
### Verify Frame Count Relationship
```bash
# face ≤ pose ≤ appearance
FACE_COUNT=$(cat $OUTPUT/$UUID.face.json | jq '.frames | length')
POSE_COUNT=$(cat $OUTPUT/$UUID.pose.json | jq '.frames | length')
APP_COUNT=$(cat $OUTPUT/$UUID.appearance.json | jq '.frames | length')
echo "Face: $FACE_COUNT, Pose: $POSE_COUNT, Appearance: $APP_COUNT"
# Expected: Face ≤ Pose ≤ Appearance
```
---
*Document Version: 1.0*
*Last Updated: 2026-07-19*