AI Interfaces
Token Stream Timeline
A delivery trace for first-token latency, reasoning time, streaming speed, and total completion.
Installation
Copy and paste the code into your project.
Accessibility notes
Every phase exposes its timing as text; decorative duration rails are hidden from assistive technology.
Token stream
Response delivery trace
- First token240ms
- Reasoning680ms
- Streaming1.4s
Preview accent
const phases = [{ label: "First token", value: "240ms", width: "w-[20%]" }, { label: "Reasoning", value: "680ms", width: "w-[38%]" }, { label: "Streaming", value: "1.4s", width: "w-[72%]" }];
export function TokenStreamTimeline() {
return <section className="w-full max-w-md rounded-xl border border-white/12 bg-[#0b0f14]/92 p-4 shadow-2xl">
<header className="flex items-start justify-between"><div><p className="text-xs font-semibold uppercase tracking-[0.18em] text-[#40E0D0]">Token stream</p><h3 className="mt-1 text-base font-bold text-white">Response delivery trace</h3></div><span className="rounded-full bg-[#40E0D0]/10 px-2.5 py-1 text-[0.65rem] font-bold text-[#d8fffb]">48 tok/s</span></header>
<ol className="mt-5 space-y-3">{phases.map((phase) => <li key={phase.label} className="grid grid-cols-[5rem_1fr_auto] items-center gap-3"><span className="text-xs text-slate-400">{phase.label}</span><span className="h-2 rounded-full bg-white/[0.06]"><span className={"block h-full rounded-full bg-[#40E0D0] " + phase.width} /></span><span className="text-xs font-bold text-white">{phase.value}</span></li>)}</ol>
</section>;
}<TokenStreamTimeline />Related components
Context compression
Compaction preview
- Keep
Decisions
1.8k
- Summarize
Discussion
8.4k → 1.2k
- Drop
Duplicates
2.1k
AI Interfaces
Context Compression Preview
A compaction preflight showing which conversation sections will be kept, summarized, or dropped.
Context window
25.4k / 32k
- System
- 4.2k
- Conversation
- 12.8k
- Sources
- 8.4k
6.6k tokens available for response
AI Interfaces
Context Window Meter
A token allocation breakdown for system instructions, conversation history, and retrieved sources.
Semantic cache
Response reuse
- Hit rate
- 68%
- Saved
- $42.18
- Latency
- −310ms
Similarity threshold 0.92 · Entries expire after 24 hours
AI Interfaces
Semantic Cache Monitor
A compact cache dashboard for hit rate, cost savings, latency reduction, and similarity policy.