SECUREXSECURITY ENGINEERINGNVR 文档

语义搜索

语义搜索让用户用自然语言寻找录像,例如“穿红色外套的人”“白色货车停在门口”。它不是普通关键词搜索,而是把图像与文本编码到同一个向量空间,再按相似度检索。

关键前提

用于 embeddings role 的模型必须专门为图文检索训练。普通 Chat/Description 模型即使能返回向量,也不代表这些向量适合检索;最常见的故障就是“没有报错,但搜索结果完全不相关”。

工作流程

  1. 系统对可搜索的图像/事件生成 image embedding。
  2. 用户输入自然语言,系统生成 text embedding。
  3. 向量数据库计算相似度并返回候选结果。
  4. 再结合时间、摄像头、对象和 Zone 等结构化条件过滤。

配置设计

genai:
  local_embed:
    provider: openai
    base_url: http://local-ai:8080/v1
    model: multimodal-embedding-model
    roles:
      - embeddings

为什么第一次很慢

启用或更换 embedding 模型后,历史数据可能需要重新建立索引。摄像头数量和历史事件越多,首次建立索引越久。应显示任务进度,而不是让 UI 看起来“卡死”。

搜索技巧

  • 描述视觉属性和动作:“黑色轿车驶入车道”比“我的车”更可靠。
  • 结合时间、摄像头、Zone 过滤,可以显著提高精度。
  • 模型不擅长读取很小的文字时,不要把语义搜索当 OCR/LPR 使用。

排错

  • 完全无结果:确认 embeddings provider 初始化成功和索引数量。
  • 结果随机:高度怀疑使用了不适合检索的模型。
  • 新事件可搜、旧事件不可搜:历史 re-index 尚未完成。
  • 速度慢:检查 embedding 推理设备和向量库 I/O。

Semantic search

Semantic search lets users find video with natural language such as “person in a red jacket” or “white van stopped at the entrance.” It encodes images and text into the same vector space and retrieves by similarity.

Critical requirement

The model assigned to the embeddings role must be trained for multimodal retrieval. A chat model may return vectors without error, but those vectors can produce meaningless search results.

Pipeline

  1. Create image embeddings for searchable events/images.
  2. Create a text embedding for the user query.
  3. Rank stored vectors by similarity.
  4. Apply structured filters such as time, camera, object, and zone.

Example provider

genai:
  local_embed:
    provider: openai
    base_url: http://local-ai:8080/v1
    model: multimodal-embedding-model
    roles: [embeddings]

Why initial indexing is slow

Enabling or changing the embedding model may require historical re-indexing. Large event histories can take time; the UI should expose progress rather than appearing stuck.

Search tips

  • Describe visible attributes/actions rather than personal ownership.
  • Combine semantic queries with time/camera/zone filters.
  • Do not treat semantic search as OCR/LPR for tiny text.

Troubleshooting

  • No results → provider initialization and index count.
  • Random results → wrong model type for retrieval.
  • Only new events searchable → re-index still running.
  • Slow queries → inspect embedding hardware and vector-store I/O.

Búsqueda semántica

Permite buscar vídeo con lenguaje natural, por ejemplo “persona con chaqueta roja” o “furgoneta blanca en la entrada”. Imágenes y texto se convierten al mismo espacio vectorial.

Requisito crítico

El modelo del rol embeddings debe estar entrenado para retrieval multimodal. Un modelo de chat puede devolver vectores sin error y aun así producir resultados inútiles.

Flujo

  1. Generar embeddings de imágenes/eventos.
  2. Generar embedding del texto.
  3. Ordenar por similitud.
  4. Aplicar filtros de tiempo, cámara, objeto y zona.

Provider

genai:
  local_embed:
    provider: openai
    base_url: http://local-ai:8080/v1
    model: multimodal-embedding-model
    roles: [embeddings]

Indexación inicial

Cambiar el modelo puede requerir reindexar el histórico. Debe mostrarse progreso.

Consejos

  • Describa atributos visibles y acciones.
  • Combine con filtros estructurados.
  • No sustituye OCR/LPR de texto pequeño.

Problemas

  • Sin resultados → provider/índice.
  • Resultados aleatorios → modelo inadecuado.
  • Solo eventos nuevos → reindex en curso.
  • Lento → hardware de embeddings o I/O.
输入关键词开始搜索