Stretchy technical whitepaper

This whitepaper describes how Stretchy indexes, scores and serves full-text search at the edge and on disconnected networks, reducing search infrastructure complexity and overhead.

Architecture overview

Stretchy is engineered around hybrid keyword (BM25 full-text) and semantic vector search running on Cloudflare's network. In edge environments, it pairs with Cloudflare Workers and D1 serverless storage to deliver fast, low-overhead search. For local deployments, it packages search and retrieval directly into a lightweight container with zero cluster coordination tax.

Index format and BM25 scoring

Stretchy avoids dynamic mapping exceptions by employing an 'Index Flattened, Store Original' strategy. Raw payloads are preserved exactly as ingested, while nested attributes are indexed into high-performance numeric and keyword trees. Text fields are scored using BM25 relevance with customizable term saturation (k1) and document length normalization (b), complemented by semantic vector matching.

Probabilistic anomaly detection

Unlike legacy search stacks that require separate telemetry processing clusters, Stretchy integrates Welford's algorithm directly into ingestion pipelines. Online mean and variance are computed in single-pass O(1) time and O(1) space, instantly flagging statistical outliers via dynamic Chebyshev Z-score thresholds.

Deployment: edge, on-prem and air-gapped

A single Stretchy Docker container requires under 150MB of RAM, making it deployable on edge gateways, developer laptops, and sovereign on-premise servers. Rich document ingestion for PDF, DOCX, and XLSX is powered by lightweight native parsers, eliminating the 100MB+ memory tax of external tools.

Security and licensing

Enterprise editions feature air-gapped Ed25519 signature verification, MAC/UUID hardware node-locking, and encrypted monotonic clock-rollback defenses, ensuring sovereign operation without cloud dependencies.