Sub-millisecond vector similarity cache for OpenAI, Claude, and Gemini. Cut inference costs by 30-60%.
POST /v1/cache/check (Similarity Match)
POST /v1/cache/set (Store Prompt/Response)
POST /v1/cache/batch-check (Batch Lookups)
GET /v1/cache/stats (Live Savings Metrics)
GET /v1/health (Healthcheck)