Media · Digital library

Agent search over 100 million documents

Vertex AI Search deployment for a subscription reading platform: English and Spanish data stores, indexing pipelines, observability, and intent-aware ranking.

A digital library with roughly a hundred million documents needed search that understood what a reader was trying to do — find a specific title, browse a topic, locate a form — rather than matching keywords. They also needed the index to stay honest as documents changed and were removed.

Phase one: full deployment

  • Data-store design and bulk import from the client's Databricks environment, with a retry mechanism for the transient backpressure that Vertex AI Search returns on large imports (it is not a terminal error, and treating it as one wastes days).
  • Metadata update and document deletion pipelines so the index reflects the catalogue.
  • An observability suite: Cloud Logging routed into BigQuery telemetry tables so the client can see query volume, latency and zero-result rates without asking us.

Phase two: intent-aware ranking

The second engagement adds English and Spanish multi-data-store search, automates Discovery Engine CRUD, and applies ranking formulas per user intent — four intents, each with its own boost specification — followed by a stress test and user acceptance. The client keeps ownership of the intent classifier and gateway; we own the search platform behind it.

Search is a platform, not a project. The observability pipeline was in the first statement of work for a reason: you cannot tune ranking you cannot measure.

All case studies

Talk to us

Tell us what system the answer lives in and who needs it. We'll reply with a view on whether it's a two-week assessment, a five-week pilot, or something else.

akash@insightnext.tech

InsightNext on LinkedIn