Back to Blog Engineering

How Velocity Scores Contextual Relevance: A Technical Overview

Marcus Webb ·
Abstract technical visualization of contextual relevance scoring

The core job of Velocity's matching engine is straightforward to state and genuinely difficult to execute: given a conversation turn, find the most relevant advertiser category in under 50 milliseconds at the 95th percentile. Everything else in the pipeline depends on that decision landing correctly.

What we mean by "contextual relevance" is specific. We are not scoring whether a user is likely to click. We are scoring whether an advertiser's topic maps cleanly onto what the user is currently asking about and what the assistant is responding with. Click prediction is a second-order concern. First-order is semantic match.

How a Conversation Turn Becomes a Signal

When a publisher's chat application sends us a conversation turn via the Velocity SDK, we receive a structured payload containing the user message, the assistant's response in progress (partial or full), and metadata like session position and turn count.

The first processing step is context windowing. We do not score only the current turn in isolation; we consider the last three turns, up to 1,200 tokens, as a sliding context window. This matters because users rarely front-load full intent in a single message. A question like "which one would you recommend?" is unscoreable without knowing that the previous two turns were about comparing project management tools.

From this window, we run a lightweight embedding model that converts the text into a dense vector representation. We keep this model small by design: 384 dimensions, under 80ms inference time including tokenization, running on-instance without external API calls. The embedding captures semantic meaning without tracking user identity or storing session history beyond the active request.

The Scoring Pipeline

The scoring pipeline runs three sequential checks before producing an output score.

First, topic detection. We maintain a taxonomy of roughly 600 advertiser-relevant topic categories, organized in a two-level hierarchy (for example, Home Improvement contains Power Tools as a subcategory). The conversation embedding is compared via cosine similarity against pre-computed centroid vectors for each leaf node. Topics with similarity above a configurable threshold (default 0.72) advance to the next stage.

Second, intent classification. Not every conversation mentioning a product category signals purchase intent. "My dishwasher broke last year" and "I'm trying to decide which dishwasher to buy" both score high on the Home Appliances topic, but only one carries commercial intent. We run a lightweight binary classifier over the top-ranked topic candidates that outputs an intent probability. Candidates below 0.4 probability are dropped. This is one of the places where we have the most room to improve, and we are actively collecting labeled data to refine it.

Third, publisher filter application. Before any score is returned to the matching stage, we apply the publisher's configured category exclusions. If a publisher has excluded Finance and Gambling categories from their ad slot, any candidate in those categories exits here regardless of relevance score. This layer is pure lookup and adds under 1ms. The full range of publisher-configurable safety and exclusion controls is described in the brand suitability guide.

Why Latency Is a First-Class Concern

The scoring step must complete before the LLM response stream begins to flow to the user. In practice, for streaming responses, we have a pre-stream window of roughly 60-80ms between when the LLM starts generating and when the first tokens arrive at the publisher's rendering layer.

That gives us a hard deadline. If the relevance score is not ready before that window closes, we have two options: hold the stream (visible latency increase for the user) or skip the ad for that turn (missed revenue for the publisher). We chose to optimize aggressively for staying within the window rather than occasionally overshooting and holding.

Our current p95 end-to-end latency for the scoring pipeline is 41ms, measured from SDK payload receipt to scored match returned. We target 50ms as our SLA boundary. The 9ms margin is not comfortable, but it is real under production load. We have spent significant time on the embedding inference path to get there, including quantizing the model weights to INT8.

What Gets Scored and What Does Not

One design decision we made early was to score the conversation pair (user query plus partial assistant response) rather than the query alone. This turns out to matter: the assistant's response is often more topic-dense than the user's question, which tends to be telegraphic. Scoring both increases topic signal without requiring the user to be explicit.

We do not score user-injected content outside the structured turn format, and we do not attempt to score turns where the conversation context is flagged for content safety exclusion (handled by a separate pipeline). If the context window contains categories marked as unsafe, the relevance scorer receives a suppressed input and returns a no-match result without logging the underlying content.

Limitations We Have Learned to Work With

The cosine similarity approach works well for clear topic matches and breaks down for ambiguous or multi-topic conversations. A turn about sustainable investing while planning a home renovation can confuse the classifier, which tends to pick the higher-similarity topic and discard the other rather than scoring both. We handle this by keeping a fallback pass that returns two candidates when both exceed a lower similarity threshold (0.60), but the advertiser fill logic only serves one format per turn.

We are also aware that the 600-topic taxonomy is incomplete. There are product categories that appear regularly in AI chat conversations that do not yet map cleanly to advertiser demand. Pet insurance is a good example: taxonomically adjacent to both General Insurance and Pet Care, with different advertiser types buying against each. We resolve ambiguities like this through manual taxonomy review cycles every six weeks.

We are not saying our current approach is optimal. It is the approach that meets the latency budget, works at our current scale, and produces match quality we can measure and improve. The taxonomy will grow. The intent classifier will get better labeled data. The threshold calibration is something we revisit quarterly as advertiser demand patterns shift.

What the Score Actually Drives

The final output of the scoring pipeline is a topic ID, an intent probability, and a composite relevance score from 0.0 to 1.0. The matching engine uses the relevance score as a filter floor: matches below 0.50 composite score do not enter the auction. Above that floor, the auction proceeds using standard CPM bid logic, and the relevance score is not used to adjust bids, only to qualify or disqualify candidates.

Publishers can see aggregate relevance score distributions for their topic mix in the Studio and Platform dashboards. This matters because a publisher whose conversation topics cluster around categories with thin advertiser demand will see low fill rates regardless of relevance score accuracy. Understanding topic distribution helps publishers set realistic fill rate expectations and helps us identify where advertiser supply gaps are.

The scoring system is a cascade of lightweight models with hard time budgets, a manually maintained taxonomy, and a set of thresholds that we adjust as we gather data. The constraints are real and the tradeoffs are intentional. If you have questions about how your specific topic coverage maps to the scoring pipeline, the technical integration docs in the Velocity dashboard go deeper on taxonomy browsing and per-category relevance diagnostics.

More from the blog
Abstract visual representing content safety filtering in AI responses
Publisher Guide
Content Safety and Brand Suitability in AI Responses
Priya Nair ·
Abstract representation of SDK integration into a chat platform
Tutorial
Integrating Velocity into Your Chat Application: A Step-by-Step Guide
Marcus Webb ·

Start earning from your AI assistant conversations.