This guide walks through a complete Velocity SDK integration for a Node.js chat application that uses streaming responses from a major LLM provider. The steps translate directly to Python and other runtimes; the concepts are the same regardless of language.
Before starting, you will need a Velocity publisher account with an API key. If you have not yet applied for publisher access, the registration flow is at veiocity.com/login/register.html. Keys are provisioned within one business day for approved publishers.
Step 1: Install the SDK
The Velocity SDK is available via npm for Node.js applications:
npm install @velocity-ads/sdk
For Python applications:
pip install velocity-ads-sdk
The SDK has no runtime dependencies beyond a standard HTTP client. The package size is under 8KB gzip, which we designed deliberately to keep it viable as a server-side dependency in latency-sensitive applications. You initialize it once at application startup with your API key and publisher configuration.
const { VelocityClient } = require('@velocity-ads/sdk');
const velocity = new VelocityClient({
apiKey: process.env.VELOCITY_API_KEY,
publisherId: 'your-publisher-id',
adSlotConfig: {
maxAdsPerSession: 3,
minTurnsBetweenAds: 2
}
});
The adSlotConfig block controls frequency. maxAdsPerSession caps the number of ad insertions across an entire conversation session. minTurnsBetweenAds sets the minimum number of turns that must pass between consecutive ads. Both values can be adjusted in the dashboard, and dashboard values will override SDK initialization values if you prefer to manage frequency centrally.
Step 2: Hook Into the Turn Lifecycle
Velocity needs to see the conversation context before the LLM response is streamed to the user. The SDK exposes a scoreConversationTurn method that takes the current conversation context and returns a match result asynchronously.
async function handleChatTurn(session, userMessage) {
const context = {
sessionId: session.id,
turnIndex: session.turns.length,
conversationHistory: session.turns.slice(-3),
currentUserMessage: userMessage
};
const matchResult = await velocity.scoreConversationTurn(context);
const llmStream = await openai.chat.completions.create({
model: 'gpt-4o',
messages: buildMessages(session.turns, userMessage),
stream: true
});
return { matchResult, llmStream };
}
The call to scoreConversationTurn fires before the LLM call in this example. In practice, you can fire both requests in parallel using Promise.all, because the Velocity scoring call is typically faster than the LLM call's first response token. The key constraint is that the match result must be available before you start rendering the ad into the response stream.
Step 3: Inject the Ad Into the Stream
When matchResult.hasAd is true, the match result contains a fully rendered ad card as an HTML snippet. The SDK handles the rendering; you handle where it appears in the stream.
The recommended insertion point is after a natural paragraph break in the assistant's response, specifically at the point where the response transitions from the first major thought to a supporting or follow-up point. The SDK provides a helper, findStreamInsertionPoint, that watches the token stream and signals when a paragraph boundary is reached.
const renderer = velocity.createStreamRenderer({
matchResult,
insertionStrategy: 'first-paragraph-break'
});
for await (const chunk of llmStream) {
const token = chunk.choices[0]?.delta?.content || '';
const output = renderer.processToken(token);
if (output.adCard) {
yield output.adCard;
}
if (output.token) {
yield output.token;
}
}
The adCard value is an HTML block containing the sponsored card markup. It always includes the Velocity attribution badge ("Sponsored") as a required element. Do not strip or hide this element. FTC guidelines on sponsored content require clear labeling, and publisher accounts are suspended for ad rendering that removes or obscures the sponsorship attribution.
Step 4: Configure Category Exclusions
Before going live, set your category exclusion preferences in the Velocity dashboard under Settings > Ad Preferences. This step is easy to skip and worth doing carefully. A full explanation of how exclusions interact with the safety layer is in the content safety and brand suitability guide. The defaults allow all advertiser categories that pass the platform-level safety filters, which may be broader than what fits your assistant's use case.
If your assistant is oriented around a specific domain, like personal finance, productivity tools, or health information, think about which adjacent categories your users would find jarring to see advertised. Finance publishers often exclude Gambling and Crypto. Health-oriented assistants frequently exclude Tobacco and Alcohol. These are not universal rules; they depend on your product and your reader relationship.
Category exclusions can be updated at any time through the dashboard and take effect immediately for new sessions. They do not require an SDK update or redeployment.
Step 5: Test in Sandbox Mode
The SDK includes a sandbox mode that returns synthetic ad cards against real conversation scoring without billing the account or serving real advertiser creative. Activate it by passing sandbox: true to the client initialization or setting the VELOCITY_SANDBOX=true environment variable.
In sandbox mode, the scoring pipeline runs fully, the match result reflects real topic classification against your test conversations, and the rendered ad card is clearly marked as a test ad. This lets you verify that the rendering, stream injection timing, and frequency controls work as expected before exposing the integration to users.
A few things worth testing explicitly in sandbox: long conversations where the session ad cap should kick in and prevent additional insertions; turns in excluded categories where no ad should appear; and very short turns where the scoring engine returns no match. All three should produce clean no-ad output with the response stream flowing normally.
Step 6: Go Live and Monitor
Remove the sandbox flag, deploy, and watch the dashboard. The Publisher Overview shows real-time turn volume, match rate, fill rate, and estimated RPM. For a new integration, expect fill rate to start low in the first few hours as the session matching warms up, and normalize toward your expected range within the first full day of traffic.
The most common first-day issue we see is incorrect session ID handling, where the publisher is generating new session IDs on each turn rather than maintaining a consistent ID across a session. The frequency controls depend on session continuity. If session IDs reset every turn, the maxAdsPerSession limit never accumulates correctly, and users can see too many ads. Check the Sessions table in the dashboard to confirm sessions are being tracked with continuity.
One thing we are direct about with publishers: the first week of metrics is not fully representative of steady-state performance. Fill rates stabilize as the contextual matching system learns your traffic's topic distribution. If fill rate looks low in the first week, look at the Topic Mix report in the dashboard before drawing conclusions. Thin advertiser demand in your primary topic categories is the most common cause of fill rate underperformance, and it tells you something about advertiser supply gaps rather than integration quality.
The full API reference and SDK documentation live in your publisher dashboard under Developer Docs. We keep the changelog up to date and version the SDK semantically, so you can pin to a minor version and not worry about breaking changes during a business-critical period.