Local-First Semantic Search on Windows: Running SigLIP2 and DINOv3 with Tauri 2 and ONNX
I built PixaFind for the search you do when you remember a file's contents but have forgotten its name. You might remember a cat sitting in sunlight, a whiteboard from a meeting, or a paragraph about a budget. None of those memories give you a path to type into Explorer. On Windows, I use a Tauri 2 shell, a Rust backend, ONNX Runtime inference, and local SQLite vector storage. SigLIP2 handles text-to-image search. DINOv3 handles image-to-image similarity. For documents, I combine dense embeddings with a BM25 keyword index. The engineering cost sits in the indexing pipeline: decoding files, preparing model inputs, running inference, and keeping the index consistent as users edit their folders. Put the inference boundary on the user's machine I keep file access, inference, and retrieval in the Rust backend. Users interact with a floating search window through a global hotkey, Ctrl+Space , without switching away from their current application. The native Tauri shell weighs about 5 MB. That figure describes the executable, not the full installation. Users also need model weights, ONNX Runtime dependencies, and index storage. The release README describes a one-time model download of roughly 500 MB. Tauri uses the system webview on Windows. I avoid shipping a separate Chromium distribution, but I still have to account for webview memory, model memory, and GPU allocations. A small executable does not imply a 5 MB runtime footprint. This diagram shows the architecture at the component level, rather than private source-code interfaces: flowchart LR F["Selected folders"] --> W["Filesystem watcher"] W --> C["Rust: mtime and SHA-256 change detection"] C --> P["Decode images or extract document text"] P --> O["ONNX Runtime: DirectML, CUDA, or CPU"] O --> V["Vector DB: local SQLite storage"] P --> B["BM25 keyword index"] V --> R["Rust retrieval and ranking"] B --> R H["Global hotkey"] --> U["Spotlight UI: Tauri 2"] U --> Q["Query or selected image"] Q --> O R --> U I separate background indexing from interactive retrieval because the two workloads compete for resources. During indexing, a user may still want to search the files that have finished processing. An implementation needs bounded work queues and an explicit policy for sharing GPU time; launching an inference task for every filesystem event would exhaust memory before it improved throughput. Use SigLIP2 for descriptions, DINOv3 for visual neighbors For a query such as cat in a sunbeam , I need an image representation that shares a space with a text representation. SigLIP2 supplies that alignment. During indexing, I encode an image. At search time, I encode the user's description and compare it with the stored image embeddings. The user does not need to tag the image or include the word cat in its filename. Zero-shot matching means I can search with a description without training a classifier for each phrase. For normalized embeddings, cosine similarity reduces to a dot product: query text -> SigLIP2 text encoder -> q photo -> SigLIP2 image encoder -> v score(q, v) = sum(q[i] * v[i]) A description such as blurry photo of whiteboard asks for visual meaning. It does not guarantee that I can retrieve a sentence written on that whiteboard. Users still depend on the image resolution and the model's ability to recognize the scene. Exact text inside an image requires an OCR path; semantic image matching alone does not establish one. When a user selects a photo and asks for similar images, I use DINOv3. With that route, I compare visual embeddings from the selected image against visual embeddings from the library. The user can look for similar compositions or objects without translating them into a sentence. I keep the two embedding spaces separate. A DINOv3 vector and a SigLIP2 text vector do not acquire compatible meaning because both contain floating-point numbers. Each search route needs the matching model, preprocessing, and vector collection. flowchart TD T["Text: cat in a sunbeam"] --> ST["SigLIP2 text encoder"] I["Indexed image"] --> SI["SigLIP2 image encoder"] ST --> S["Compare within SigLIP2 space"] SI --> S X["Selected image"] --> DX["DINOv3 encoder"] I --> DI["DINOv3 encoder"] DX --> D["Compare within DINOv3 space"] DI --> D S --> A["Rank image results"] D --> A Run ONNX inference without guessing the model contract I use ONNX Runtime to execute the exported models on Windows. PixaFind supports DirectML and CUDA acceleration, with a CPU fallback. DirectML covers compatible DirectX 12 GPUs across vendors; CUDA targets supported NVIDIA hardware. I still need to ship a runtime that contains the selected execution provider. Enabling a Rust Cargo feature does not install CUDA libraries or give an ONNX Runtime binary DirectML support. I also need to check the exported graph's operators, tensor shapes, and provider compatibility. The following example uses the documented API from ort 2.0.0-rc.10. It illustrates a DirectML image-embedding inference boundary, not a copy of PixaFind's private implementation. The caller supplies the actual model path, tensor names, spatial dimensions, and preprocessed pixels from its export contract. use std::path::Path; use ort::{ execution_providers::DirectMLExecutionProvider, session::Session, value::TensorRef, }; fn open_image_encoder(model_path: &Path) -> ort::Result { Session::builder()? .with_parallel_execution(false)? .with_memory_pattern(false)? .with_execution_providers([ DirectMLExecutionProvider::default().build(), ])? .commit_from_file(model_path) } fn infer_embedding( session: &mut Session, input_name: &str, output_name: &str, height: usize, width: usize, pixels_nchw: &[f32], destination: &mut Vec , ) -> ort::Result { // The tensor view borrows the caller's preprocessed image buffer. let input = TensorRef::from_array_view(( [1usize, 3, height, width], pixels_nchw, ))?; let outputs = session.run(ort::inputs![input_name => input])?; let (_shape, embedding) = outputs[output_name].try_extract_tensor:: ()?; // Preserve the result after the runtime output values leave scope. destination.clear(); destination.extend_from_slice(embedding); Ok(()) } I reuse the session across images instead of loading model weights for each forward pass. I borrow the input buffer and copy the output into caller-owned storage because I need that embedding after the runtime releases its output values. The caller can reuse the destination allocation across files. The caller must decode the image, apply the export's resize or crop rule, normalize channels with the correct constants, and produce the expected NCHW layout. It must also verify that the chosen output contains a pooled image embedding rather than logits or a sequence of patch features. Changing any of those choices can change retrieval quality without triggering a tensor-type error. For SigLIP2 text inference, I need a separate input contract: the matching tokenizer, token IDs, padding policy, and any required mask tensors. I cannot substitute image preprocessing or an unrelated tokenizer. Microsoft's DirectML execution-provider documentation requires sequential graph execution and disabled memory-pattern optimization. It also prohibits concurrent Run calls on the same DirectML session. That constraint affects worker design: I serialize access to a shared session, or provision separate sessions and accept their additional memory cost. The documentation now describes DirectML as being in sustained engineering, with new Windows deployment development moving toward WinML. For developers choosing a runtime today, that distinction matters even when an existing DirectML deployment continues to work. Keep exact document terms alongside dense embeddings For document search, I combine dense vector retrieval with BM25. The release README documents approximately 30 document extensions, including .pdf , .docx , .pptx , .md , .txt , and source-code formats. The product website advertises a broader 76-plus-format count. I use the release README's narrower document scope here rather than treating those two counts as interchangeable. A developer searching for vacation policy may want a paragraph about annual leave. Dense embeddings can retrieve that semantic relationship. Another developer searching for ERR_CONNECTION_RESET , a function name, or an invoice identifier needs exact terms. BM25 covers that lexical route. For a hybrid implementation, I would use this retrieval sequence: Extracted document text | +-> bounded text chunks -> dense embeddings -> semantic candidates | +-> tokenized text -> BM25 index -> keyword candidates | merge, deduplicate, rank by file This sketch describes a design pattern, not a claim about PixaFind's private chunk size or fusion formula. The public release repository documents hybrid search without publishing those internals. Developers need a ranking policy because BM25 scores and vector similarities have different scales. Adding the raw values can let one route dominate by numerical range. Rank-based fusion offers one option: assign a contribution according to a candidate's position in each list, then deduplicate document chunks before presenting file results. Tuning that policy requires queries from the target library, including both paraphrases and exact identifiers. Text extraction also sets a ceiling on recall. A scanned PDF may contain pixels without a text layer. Office documents can include tables whose reading order affects the extracted text. Source-code tokenization can split identifiers that users expect to search as a unit. I would inspect extraction output before blaming the embedding model for a missing result. Avoid paying for the same forward pass twice I use modification-time checks and SHA-256 content hashing to avoid redundant inference on unchanged files. Reading metadata costs less than decoding an image and running a vision encoder. Hashing costs a full read of the file, so I do not treat it as free. The conceptual decision tree looks like this: Filesystem
Comments
No comments yet. Start the discussion.