# Search Improvements for Core Team ## Overview This document outlines the search improvements implemented in Studio that may require Core API support or coordination. --- ## 1. Keyword Search - Result Merging ### Current Implementation (Studio-side) ```typescript function mergeResults(results: any[]): any[] { // Sort by start_frame const sorted = [...results].sort((a, b) => (a.start_frame || 0) - (b.start_frame || 0)) // Merge overlapping results (gap <= 30 frames) for (const r of sorted) { const last = merged[merged.length - 1] if (last && last.file_uuid === r.file_uuid) { const gap = r.start_frame - last.end_frame if (gap <= 30) { // Merge: extend end_frame, keep best summary last.end_frame = Math.max(last.end_frame, r.end_frame) if (r.similarity > last.similarity) { last.summary = r.summary last.similarity = r.similarity } continue } } merged.push({ ...r }) } return merged.slice(0, 10) // Top 10 } ``` ### Behavior | Before | After | |--------|-------| | 20 results, overlapping segments | 10 merged results | | Same content appears multiple times | Deduplicated by frame proximity | | Short segments (1-2 seconds) | Merged into longer segments | ### Core API Considerations **Option A: Server-side Merging** - Add `merge=true` parameter to `/api/v1/search/smart` - Server returns pre-merged results - More efficient, less data transfer **Option B: Client-side (Current)** - Studio fetches 30 results, merges locally - More flexible, no Core API changes needed --- ## 2. Minimum Duration Enforcement ### Problem Some search results are too short (1-2 seconds) to provide meaningful context. ### Proposed Solution Add `min_duration` parameter to search endpoints: ```bash POST /api/v1/search/smart { "query": "Audrey Hepburn", "min_duration": 10, // Minimum 10 seconds "limit": 10 } ``` ### Core API Impact | Endpoint | Change | |----------|--------| | `/api/v1/search/smart` | Add `min_duration` param | | `/api/v1/search/llm-smart` | Add `min_duration` param | **Implementation:** 1. Filter results where `end_time - start_time < min_duration` 2. Or expand short results by including adjacent chunks --- ## 3. Frame-based Range Adjustment ### Current Studio Implementation Users can adjust playback range by frame number: ```vue
``` ### Core API Requirements No changes needed - Studio uses existing `start_frame`/`end_frame` fields. **Display format:** ``` F100–F350 (4:10–14:35) ``` --- ## 4. Mark & Export Feature ### Studio Implementation ```typescript const markedResults = ref([]) function toggleMark(r: any) { const key = `${r.file_uuid}:${r.start_frame}-${r.end_frame}` const idx = markedResults.value.findIndex(m => `${m.file_uuid}:${m.start_frame}-${m.end_frame}` === key ) if (idx >= 0) { markedResults.value.splice(idx, 1) } else { markedResults.value.push({...r}) } } function exportMarked() { const data = { exportedAt: new Date().toISOString(), count: markedResults.value.length, results: markedResults.value } // Download as JSON } ``` ### Export Format ```json { "exportedAt": "2026-07-20T12:00:00.000Z", "count": 3, "results": [ { "file_uuid": "abc123", "file_name": "Roman Holiday.mp4", "start_frame": 100, "end_frame": 350, "start_time": 4.16, "end_time": 14.58, "summary": "Audrey Hepburn speaking...", "similarity": 0.85 } ] } ``` ### Future Core API Integration **Possible endpoints:** ``` POST /api/v1/marks # Save marked segment GET /api/v1/marks # List saved marks POST /api/v1/marks/export # Export marks as EDL/JSON ``` --- ## 5. Vector/Keyword Weight Ratio ### Current State `/api/v1/search/llm-smart` uses RRF (Reciprocal Rank Fusion): - Vector search: 50% - Keyword (BM25): 50% ### Proposed Change Increase vector weight for semantic relevance: ```python # Current score = 0.5 / (k + rank_vector) + 0.5 / (k + rank_keyword) # Proposed VECTOR_WEIGHT = 0.7 KEYWORD_WEIGHT = 0.3 score = VECTOR_WEIGHT / (k + rank_vector) + KEYWORD_WEIGHT / (k + rank_keyword) ``` ### Core API Change Required Add `vector_weight` parameter: ```bash POST /api/v1/search/llm-smart { "query": "two people talking", "vector_weight": 0.7, # Default: 0.5 "limit": 10 } ``` --- ## 6. Agent Search - Question Templates ### Studio Implementation Pre-defined question templates shown below search input: ```typescript const QUESTION_TEMPLATES = [ { category: '找檔案', examples: [ '找出包含 {人物} 的影片', '找出 {年份} 年的影片' ] }, { category: '找人物', examples: [ '{人物} 出現在哪些影片?', '{人物} 和 {人物} 第一次同框' ] }, { category: '找內容', examples: [ '找出討論 {主題} 的片段', '找出 {時間點} 發生什麼事' ] } ] ``` ### Core API Requirements No changes needed - templates are Studio-side UI. --- ## 7. VLM Search Integration (Future) ### Proposal Add visual-language model search capability: ```bash POST /api/v1/search/vlm { "query": "person wearing red dress", "file_uuid": "optional" } ``` ### Response ```json { "results": [ { "file_uuid": "abc123", "start_frame": 1000, "end_frame": 1050, "description": "Woman in red dress walking", "confidence": 0.92 } ] } ``` ### Core Team Considerations 1. Model selection (CLIP, BLIP, etc.) 2. GPU requirements 3. Indexing strategy 4. Latency expectations --- ## Summary of Core API Changes | Feature | Priority | Core API Changes | |---------|----------|------------------| | Result merging | Low | Optional: `merge=true` param | | Min duration | Medium | Add `min_duration` param | | Frame adjustment | None | No changes needed | | Mark/Export | Low | Future: marks endpoints | | Vector weight | Medium | Add `vector_weight` param | | Question templates | None | Studio-side only | | VLM search | Future | New endpoint | --- ## Testing Checklist for Core Team ### 1. Search Endpoint Testing ```bash # Test basic search curl -X POST http://localhost:3002/api/v1/search/smart \ -H "X-API-Key: $KEY" \ -H "Content-Type: application/json" \ -d '{"query": "test", "limit": 30}' # Test with min_duration (if implemented) curl -X POST http://localhost:3002/api/v1/search/smart \ -H "X-API-Key: $KEY" \ -H "Content-Type: application/json" \ -d '{"query": "test", "min_duration": 10}' # Test vector weight (if implemented) curl -X POST http://localhost:3002/api/v1/search/llm-smart \ -H "X-API-Key: $KEY" \ -H "Content-Type: application/json" \ -d '{"query": "test", "vector_weight": 0.7}' ``` ### 2. Performance Testing | Metric | Target | |--------|--------| | Search latency | < 500ms | | Result count | Stable at limit=30 | | Memory usage | No significant increase | ### 3. Integration Testing - [ ] Search with file_uuid filter - [ ] Search without file_uuid (global) - [ ] Verify start_frame/end_frame accuracy - [ ] Verify time/frame consistency (fps) --- ## Questions for Core Team 1. Should result merging be server-side or client-side? 2. What's the acceptable latency for search? 3. Is `min_duration` parameter feasible? 4. Can `vector_weight` be made configurable? 5. Timeline for VLM search integration? --- *Document created: 2026-07-20* *Author: Studio Team*