Status: Proposed WICG Incubator Explainer
Author: Hemanth HM
Target Working Groups: WICG / W3C Timed Text Working Group (TTWG) / Web Real-Time Communications & Media
Web applications increasingly require sub-word, syllable-level, and continuous time-synchronized text synchronization with both audio and video media streams. Key use cases include real-time music karaoke, video closed-captioning with word tracking for cognitive accessibility, educational video lectures with synchronized text, language learning/dubbing, speech therapy, and teleprompter displays.
Today, the web platform lacks a native, declarative primitive for progressive, sub-word media alignment across HTMLMediaElement (<audio> and <video>). Web developers must resort to 60fps JavaScript requestAnimationFrame loops that query stepped, jittery HTMLMediaElement.currentTime values, causing frame drops, clock drift, and significant main-thread overhead. Furthermore, existing standards like WebVTT and TTML fail on agglutinative, inflected, and compounding languages (such as Sanskrit, German, Finnish, and Arabic) where acoustic tokens diverge from semantic words due to morphological sandhi splits.
This document proposes a unified Web platform solution for both audio and video:
- Extended Timed Text Track Cues (
VTTWordCue/VTTSyllableCue) with morphological compound split mapping for<audio>and<video>. - CSS Media Timeline Bindings (
::cue-word,::cue-syllable, and hardware-clocked CSS progressive fill sweeps). - High-Resolution
MediaTimelineAPI to drive smooth, jank-free visual synchronization directly offHTMLMediaElementaudio/video presentation timestamps (PTS).
The standard timeupdate event on <audio> and <video> elements fires at an implementation-dependent frequency (typically between 4Hz and 15Hz). Because browser media pipelines decode audio on dedicated low-level threads while dispatching events to the main JavaScript event loop, media.currentTime appears as coarse, erratic steps.
When developers attempt to build smooth word-sweep animations using JavaScript requestAnimationFrame (RAF) polling, three problems arise:
- Audio Clock vs. Display Refresh Drift: The display's V-Sync clock (e.g. 60Hz/120Hz) and the audio Digital-to-Analog Converter (DAC) sample clock drift over time without a hardware-synchronized phase lock.
- Main Thread Contention: Any heavy DOM manipulation, layout calculation, or network I/O stutters the animation loop, breaking lyrical alignment.
- Battery & CPU Drain: Running continuous un-throttled JS render loops just to interpolate a gradient position wastes battery on mobile devices.
WebVTT supports basic intra-cue timestamps (<00:01.200>word), but:
- The
::cueCSS pseudo-element only styles entire cue blocks. There is no standard selector to style active, past, or upcoming words or syllables. - There is no native mechanism for progressive spatial fill (e.g. coloring a syllable from left to right as the singer holds a vowel).
- WebVTT treats tokens as raw whitespace-separated characters without structural awareness of morphemes, phonetic duration, or multi-word compound splits.
In inflected and agglutinative languages:
- Sanskrit Sandhi: Two semantic names like kṣetrajñaḥ (क्षेत्रज्ञः) and akṣaraḥ (अक्षरः) fuse phonetically into a single chanted token kṣetrajño'kṣara (क्षेत्रज्ञोऽक्षर).
- German Compounds: Donaudampfschiffahrtsgesellschaftskapitän is a single acoustic token with multiple constituent lemmas.
- Japanese / Chinese: Lack whitespace delimiters altogether, requiring morphological token graphs rather than naive whitespace splitting.
Current web primitives force developers to pick between displaying broken, ungrammatical splits or losing time synchronization entirely.
┌─────────────────────────────────────────────────────────────┐
│ Audio Hardware Clock │
└──────────────────────────────┬──────────────────────────────┘
│
┌───────────────▼───────────────┐
│ MediaAnimationTimeline │
└───────────────┬───────────────┘
│
┌──────────────────────┴──────────────────────┐
▼ ▼
┌─────────────────────────────┐ ┌─────────────────────────────┐
│ CSS Hardware Paint │ │ Declarative Cue Engine │
│ ::cue-word(:active) { │ │ <track kind="lyrics"> │
│ --media-progress: auto; │ │ VTTWordCue / Morph-tokens │
│ } │ └─────────────────────────────┘
└─────────────────────────────┘
We propose extending CSS Animations and Pseudo-elements to bind directly to media element playback positions without JavaScript intervention:
/* Style words dynamically based on media playback status */
::cue-word(:past) {
color: var(--text-muted);
opacity: 0.6;
}
::cue-word(:active) {
color: var(--gold-primary);
/* Hardware-accelerated sweep driven by media clock */
background: linear-gradient(
to right,
var(--gold-active) 0%,
var(--gold-active) var(--cue-word-progress),
var(--text-color) var(--cue-word-progress),
var(--text-color) 100%
);
-webkit-background-clip: text;
}
::cue-word(:future) {
color: var(--text-normal);
}--cue-word-progress: Float between0.0and1.0representing continuous position through the active word.--cue-syllable-progress: Float between0.0and1.0for sub-word units.
Introduce kind="lyrics" or kind="sync-text" to the <track> element, supporting high-resolution JSON or enhanced WebVTT cues:
<!-- Audio Player (Karaoke / Chanting / Podcasts) -->
<audio id="audio-player" src="audio.mp3">
<track kind="lyrics" src="lyrics.vtt" srclang="sa" default>
</audio>
<!-- Video Player (Music Videos / Lectures / Accessible Captions / ADR Dubbing) -->
<video id="video-player" src="video.mp4" controls>
<track kind="lyrics" src="video-lyrics.vtt" srclang="en" default>
<track kind="captions" src="subtitles-word-level.vtt" srclang="en">
</video>WEBVTT
Kind: lyrics
00:00:10.000 --> 00:00:18.800
{
"tokens": [
{
"text": "पूतात्मा",
"iast": "pūtātmā",
"start": 10.000,
"end": 10.820,
"id": 10
},
{
"text": "परमात्मा",
"iast": "paramātmā",
"start": 10.820,
"end": 11.640,
"id": 11
},
{
"text": "च",
"iast": "ca",
"type": "particle",
"start": 11.640,
"end": 11.800
},
{
"text": "मुक्तानां परमा गतिः",
"iast": "muktānāṁ paramā gatiḥ",
"start": 11.800,
"end": 14.800,
"id": 12,
"sub_tokens": [
{ "text": "मुक्तानाम्", "start": 11.800, "end": 12.800 },
{ "text": "परमा", "start": 12.800, "end": 13.800 },
{ "text": "गतिः", "start": 13.800, "end": 14.800 }
]
}
]
}To allow developers using Web Animations API (WAAPI) or Canvas/WebGL to stay in phase with media playback without RAF drift:
[Exposed=Window]
interface MediaAnimationTimeline : AnimationTimeline {
constructor(HTMLMediaElement mediaElement);
readonly attribute HTMLMediaElement mediaElement;
readonly attribute double? preciseCurrentTime;
};const audioElement = document.querySelector('audio');
const timeline = new MediaAnimationTimeline(audioElement);
const wordElement = document.querySelector('#word-10');
wordElement.animate(
[
{ backgroundPosition: '0% 0%' },
{ backgroundPosition: '100% 0%' }
],
{
timeline: timeline,
timeRange: { start: 10000, end: 10820 }, // 10.00s to 10.82s
fill: 'both'
}
);- Zero Main-Thread Jitter: Syllable progress rendering is computed on the compositor thread directly off audio timebase clocks.
- First-Class Internationalization: Solves morphological compounding and sandhi across world languages.
- Accessibility by Default: Screen readers and braille displays can tap into precise live-spoken words without reverse-engineering custom JS karaoke widgets.
- Massive Energy Savings: Removes thousands of CPU wakeups per minute on mobile devices streaming synchronized audio/lyrics.
- Apple Music / Spotify Canvas / LRC Spec: Proprietary word-by-word synchronization protocols currently delivered over bespoke WebSockets / private JSON.
- SMPTE-TT / IMSC1: Rich timed text standards that describe syllable timings but lack web-native DOM/CSS integration.
- Web Audio API (
AudioContext.currentTime): Demonstrates the power of hardware-clock accuracy; this proposal bridges that precision into DOM rendering.
We encourage authors, browser engine implementers (Chromium, WebKit, Gecko), and streaming platform engineers to review this explainer, contribute edge cases, and incubate this specification within the WICG.