Back to components

AI Interfaces

Token Stream Timeline

A delivery trace for first-token latency, reasoning time, streaming speed, and total completion.

tokensstreaminglatency

Installation

Copy and paste the code into your project.

Accessibility notes

Every phase exposes its timing as text; decorative duration rails are hidden from assistive technology.

Token stream

Response delivery trace

48 tok/s
  1. First token240ms
  2. Reasoning680ms
  3. Streaming1.4s
612 output tokensComplete · 2.3s

Preview accent

TypeScript
const phases = [{ label: "First token", value: "240ms", width: "w-[20%]" }, { label: "Reasoning", value: "680ms", width: "w-[38%]" }, { label: "Streaming", value: "1.4s", width: "w-[72%]" }];

export function TokenStreamTimeline() {
  return <section className="w-full max-w-md rounded-xl border border-white/12 bg-[#0b0f14]/92 p-4 shadow-2xl">
    <header className="flex items-start justify-between"><div><p className="text-xs font-semibold uppercase tracking-[0.18em] text-[#40E0D0]">Token stream</p><h3 className="mt-1 text-base font-bold text-white">Response delivery trace</h3></div><span className="rounded-full bg-[#40E0D0]/10 px-2.5 py-1 text-[0.65rem] font-bold text-[#d8fffb]">48 tok/s</span></header>
    <ol className="mt-5 space-y-3">{phases.map((phase) => <li key={phase.label} className="grid grid-cols-[5rem_1fr_auto] items-center gap-3"><span className="text-xs text-slate-400">{phase.label}</span><span className="h-2 rounded-full bg-white/[0.06]"><span className={"block h-full rounded-full bg-[#40E0D0] " + phase.width} /></span><span className="text-xs font-bold text-white">{phase.value}</span></li>)}</ol>
  </section>;
}
Usage
<TokenStreamTimeline />

Related components

Context compression

Compaction preview

−7.5k tokens
  • Decisions

    1.8k

    Keep
  • Discussion

    8.4k → 1.2k

    Summarize
  • Duplicates

    2.1k

    Drop
Projected context5.4k / 16k

AI Interfaces

Context Compression Preview

Advanced

A compaction preflight showing which conversation sections will be kept, summarized, or dropped.

contextcompressiontokens
View component

Context window

25.4k / 32k

79% used
System
4.2k
Conversation
12.8k
Sources
8.4k

6.6k tokens available for response

AI Interfaces

Context Window Meter

Intermediate

A token allocation breakdown for system instructions, conversation history, and retrieved sources.

tokenscontextobservability
View component

Semantic cache

Response reuse

Healthy
Hit rate
68%
Saved
$42.18
Latency
−310ms
8,142 similar requests5,536 hits

Similarity threshold 0.92 · Entries expire after 24 hours

AI Interfaces

Semantic Cache Monitor

Intermediate

A compact cache dashboard for hit rate, cost savings, latency reduction, and similarity policy.

cachelatencycost
View component