AV-sync reliability audit

Gen4AVC @ ECCV 2026 · Poster

What Do Audio-Visual Synchronization Metrics Actually Measure?

Jai Kumar Sharma1,, Peeyush Tapadiya2

1 Virginia Tech  ·  2 Accenture  ·  corresponding author

The axis split

No metric is good at both

Temporal-oracle tracking versus PEAVS-proxy agreement for four AV-sync metrics Kendall tau on the temporal-shift oracle (horizontal) against Kendall tau with the PEAVS proxy (vertical). Synchformer/DeSync reaches 0.84 horizontally but only 0.07 vertically; ImageBind AV-relevance reaches 0.16 and 0.20; JavisScore 0.39 and 0.20; AV-Align 0.21 and 0.00. The upper-right region, where a metric would be strong on both axes, contains no metric. STRONG ON BOTH — EMPTY — 0.0 0.2 0.4 0.6 0.8 0.0 0.1 0.2 temporal oracle τ → PEAVS proxy τ → AV-Align — temporal oracle τ 0.21, PEAVS proxy τ 0.00 ImageBind AV-relevance — temporal oracle τ 0.16, PEAVS proxy τ 0.20 JavisScore — temporal oracle τ 0.39, PEAVS proxy τ 0.20 Synchformer/DeSync — temporal oracle τ 0.84, PEAVS proxy τ 0.07 AV-Align ImageBind JavisScore Synchformer
Synchformer/DeSync 0.84 / 0.07 JavisScore 0.39 / 0.20 ImageBind-rel. 0.16 / 0.20 AV-Align 0.21 / 0.00

Kendall τ on the temporal-shift oracle (x) against Kendall τ with the PEAVS proxy (y), 75 AVSync15 clips. ImageBind-rel. and JavisScore tie on the perceptual axis; they share an ImageBind backbone (r = 0.93) and are not independent readings.

TL;DR

Deployed AV-sync metrics measure different, incompatible reliability axes. No metric wins both temporal-offset tracking and PEAVS-proxy alignment, the four disagree with each other almost completely (Krippendorff α = 0.066), and fusing them does not help. Report AV-sync as a Reliability Card, not one bare score.

Abstract

Auditing the instruments, not the generators

Automatic AV-sync metrics are widely used to rank and train audio-visual generators, but they are rarely audited as measurement instruments. We jointly audit AV-Align, ImageBind AV-relevance, JavisScore, and Synchformer/DeSync under a common reliability protocol: controlled-distortion monotonicity, preprocessing sensitivity, rank uncertainty, cross-metric agreement, PEAVS-proxy agreement, and learned fusion.

The result is an axis split, not a single winner: Synchformer/DeSync is the strongest temporal-offset tracker (τ = 0.84), ImageBind/JavisScore better match the PEAVS human-aligned proxy (τ = 0.20) and content-disruption families, and AV-Align is the weakest standalone metric. The metrics mutually disagree (Krippendorff α = 0.066), and neither linear nor simple k-NN fusion improves PEAVS agreement over the best individual metric. We recommend reporting AV-sync as a Reliability Card (metric-family breakdowns with confidence intervals) rather than a single bare synchronization score.

Results

Six findings

Each finding is filed under the reliability axis it belongs to. Every number below is measured on AVSync15 clips under the shared protocol, with clip-bootstrap confidence intervals.

Synchformer/DeSync Oracle · temporal

The offset predictor owns temporal ordering

Synchformer/DeSync tracks a known global audio-visual shift at τ = 0.84 [0.78, 0.90] and audio speed at 0.76, against 0.16–0.39 for every other metric. It is also the only metric whose adjacent-level discriminability clears d′ ≥ 1 (d′ = 1.6 on temporal shift); everywhere else adjacent desync levels overlap by more than 16%.

ImageBind-rel. JavisScore Oracle · content

… but it does not generalize to content disruption

On fragment shuffle ImageBind leads (0.38 vs. Synchformer's 0.27) and on intermittent mute JavisScore leads (0.73 vs. 0.59). Both embedding metrics also give the highest agreement with the PEAVS human-aligned proxy, tying at τ = 0.20 while Synchformer reaches only 0.07. Competence is axis-specific: the offset predictor and the embedding metrics are good at different things.

AV-Align Preprocessing

AV-Align is the weakest and least stable metric

It is near chance on audio speed (0.04) and fragment shuffle (0.01), scores 0.00 against the PEAVS proxy, and is the only high-variance metric under should-not-change resampling: CV = 0.54 against ≤ 0.15 for the rest. Its temporal rank-flip probability is 0.66 — a coin toss on the ordering it is asked to produce. Running the official TempoTokens implementation through the same oracle reproduces every aggregate conclusion, so this is not a reimplementation artifact.

Cross-metric agreement

The metrics barely agree with one another

Treating each metric as a rater of clip sync, cross-metric agreement is Krippendorff α = 0.066 on z-scored per-clip scores — against an inter-annotator anchor of α ≈ 0.71 among PEAVS's human raters. Every pairwise Kendall τ is near zero except ImageBind-rel. vs. JavisScore (+0.78), which share a backbone and are not independent.

A PCA shows the four metrics span three orthogonal axes. A single sync number projects three dimensions onto one.

Fusion

Combining the metrics does not rescue them

A ridge meta-metric over the four scores, evaluated strictly out-of-fold with split-conformal prediction intervals, never beats the best single metric (held-out τ ≤ 0.12 vs. 0.20 against PEAVS). This is not a limitation of linearity: a leave-one-out k-NN combiner does no better (≤ 0). Naive fusion is insufficient here, though it does not rule out supervised calibration against direct human labels.

Synchformer/DeSync Ranking resolution

Close checkpoints fall inside similarity-metric noise

We generated audio for 45 AVSync15 videos with three MMAudio checkpoints and asked each metric to rank them under a paired clip-bootstrap. A far pair (small-16k vs. large-44k) is easy for everyone. On the close pair (large-44k vs. large-44k-v2) — the leaderboard regime of similar systems — only Synchformer separates the generators.

MetricFar model gapClose model gap
AV-Align≤ 0.090.35
ImageBind-rel.≤ 0.090.33
JavisScore≤ 0.090.35
Synchformer/DeSync≤ 0.090.08

Table 2 — Leaderboard risk on generated outputs. Paired clip-bootstrap flip probability (lower is better).

Re-running the oracle on the 45 generated clips reproduces the ranking (Synchformer best, τ = 0.76; AV-Align worst), so the split is not an artifact of natural clips.

Table 1

The reliability scorecard

One table, six axes, four deployed metrics. Best per column in bold; higher τ is better, lower CV and rank-flip are better. Read down the columns: the leader changes from column to column.

Metric Oracle tracking — Kendall τ Preproc. CV ↓ Rank-flip ↓ PEAVS proxy τ
Shift Speed Shuffle Mute Shift Worst fam.
AV-Align (reimpl.) 0.210.040.010.240.540.660.960.00
ImageBind-rel. 0.160.370.380.610.150.540.690.20
JavisScore 0.390.510.310.730.140.460.900.20
Synchformer/DeSync 0.840.760.270.59n/a0.191.000.07

Table 1 — Main reliability scorecard on AVSync15. Synchformer/DeSync leads temporal tracking; ImageBind and JavisScore lead PEAVS-proxy agreement (a tie); AV-Align is the weakest standalone metric. No metric wins both the temporal-oracle and the PEAVS-proxy axis. Rank-flip Shift / Worst fam. are the pure-shift and maximum-family adjacent mis-ordering probabilities. Synchformer CV is n/a (fixed ≥ 5 s input). The AV-Align row is our reimplementation; official TempoTokens scoring gives the same aggregate conclusion. Confidence intervals are in the paper's supplement.

Worst-family flip is high for every metric — but it is each metric's weak family. On the temporal-shift ranking that matters most, Synchformer flips at 0.19 (and 0.00 on speed) where the others flip at 0.46–0.66. Leaderboard gaps below a metric's minimum detectable difference (13–17% of range) are noise.

Figures

What the audit looks like

Grouped bar chart with 95% confidence intervals showing Kendall tau against known desynchronization for four AV-sync metrics across four distortion families. On temporal shift, Synchformer reaches 0.84 while JavisScore reaches 0.39, AV-Align 0.21 and ImageBind AV-relevance 0.16. On audio speed, Synchformer reaches 0.76 and JavisScore 0.51, while AV-Align is near zero at 0.04 with a confidence interval crossing zero. On fragment shuffle the ordering reverses: ImageBind leads at 0.38, JavisScore 0.31, Synchformer 0.27 and AV-Align 0.01. On intermittent mute, JavisScore leads at 0.73, ImageBind 0.61, Synchformer 0.59 and AV-Align 0.24.
Figure 1 — Agreement with controlled desynchronization (Kendall τ, 95% CI). Synchformer leads temporal and audio-speed tracking, while JavisScore and ImageBind respond more strongly to intermittent mute; AV-Align is weakest overall, with a confidence interval that crosses zero on audio speed.
Four-by-four heatmap of pairwise cross-metric rank agreement on real clips. Every off-diagonal cell is near zero: AV-Align versus ImageBind AV-relevance is -0.11, AV-Align versus JavisScore -0.07, AV-Align versus Synchformer -0.09, and Synchformer versus both embedding metrics is 0.00. The only agreeing pair is ImageBind AV-relevance versus JavisScore at 0.78, which share an ImageBind backbone. Overall Krippendorff alpha is 0.066.
Figure 2 — Cross-metric agreement (Krippendorff α = 0.066). Only the shared-backbone embedding pair agrees (τ = +0.78); Synchformer orders real clips unlike any other deployed metric.
Scatter plot of temporal-oracle Kendall tau on the horizontal axis against PEAVS-proxy Kendall tau on the vertical axis for four metrics. Synchformer sits far right but low at 0.84 and 0.07; ImageBind AV-relevance at 0.16 and 0.20 and JavisScore at 0.39 and 0.20 sit high on the perceptual axis but low on the temporal axis; AV-Align sits near the origin at 0.21 and 0.00. A shaded ideal region in the upper right, where a metric would be strong on both axes, contains no metric.
Figure 3 — The two axes are nearly orthogonal. Oracle tracking (x) against PEAVS agreement (y): detecting controlled desynchronization and matching the perceptual proxy are different problems, and no metric is high on both.
Protocol

How the audit is run

Synthetic-desync oracle
Known-magnitude desync is imposed on real, well-synced clips from four families — temporal shift, audio speed, fragment shuffle, intermittent mute — each an increasing-severity grid, so the ground-truth ordering is known by construction and no annotation is needed. We report per-clip Kendall τ between the score and the ideal “more distortion is worse” order.
Preprocessing sensitivity
Random fixed-length crops and length truncation — changes that should not alter true sync — scored by coefficient of variation, plus the bootstrap rank-flip probability for two conditions one desync step apart.
Cross-metric agreement
Each metric treated as a rater; pairwise Kendall τ and Krippendorff α on per-metric z-scored clip scores, so agreement is scale-free. PEAVS's human inter-annotator α ≈ 0.71 is the anchor.
Learned meta-metric
Ridge regression of the PEAVS score on the four metric scores, strictly out-of-fold (5-fold and leave-one-out, no leakage), with split-conformal prediction intervals; compared against every single metric.
Clips
AVSync15 (15 classes, 1,500 real highly-synced clips drawn from VGGSound). The primary audit uses 75 clips, 5 per class — the subset on which PEAVS scoring and all auxiliary analyses are aligned.
Metrics as black boxes
AV-Align reimplemented from the published TempoTokens algorithm with default peak-picking; ImageBind AV-relevance and JavisScore on the official ImageBind encoder (JavisScore: 2 s windows, 1.5 s overlap, mean of the 40% least-synced windows); Synchformer/DeSync from the released checkpoint, scored as the negative absolute expected offset. Scores are oriented so higher = better synced.
Statistics
Kendall τ-b, clip-bootstrap 95% CIs at ≥ 1k resamples, preprocessing CV, and bootstrap rank-flip — reproducible from committed code, seeds and SLURM scripts.
Robustness checks
A 2× replication (150 clips) reproduces every conclusion (Synchformer temporal τ = 0.87, α = 0.07). A second domain, 75 in-the-wild VGGSound clips, reproduces the ordering (Synchformer temporal τ = 0.63 vs. ≤ 0.27 for the rest). Two alternative Synchformer score reductions leave the temporal/perceptual split intact. Cross-metric α rises with the strength of the sync signal — 0.07 on clean clips, 0.21 on the controlled grid, 0.11 on generated clips.
The recommendation

Report a Reliability Card, not a number

The audit distills into a reusable reporting standard. An AV-sync metric should be characterized on all seven axes below — we recommend authors report this card for any metric they use to rank or train models.

Metric Reliability Card 7 axes · App. E
Reliability axisWhat it testsReport
Temporal oracleglobal-offset sensitivityτ + CI
Distortion oraclespeed / shuffle / mute sensitivityτ + CI
Preprocessing sensitivitycrop and length invarianceCV + CI
Ranking resolutionleaderboard uncertaintyflip prob.
PEAVS proxyperceptual alignmentτ + CI
Cross-metric consistencyagreement with peersrank τ / α
Direct human calibrationdirect preference checkpairwise acc.

The first two rows are the axes this audit shows are not interchangeable; the last row is the one nobody can fill in yet, and the reason the perceptual column is read as PEAVS agreement rather than perceptual ground truth.

Guidance

What to do with this today

The AV-generation community ranks — and increasingly trains — models on synchronization metrics that mutually disagree at α = 0.066, against ≈ 0.71 among PEAVS's human annotators. An unreliable metric does not just mis-rank a leaderboard; in a preference-optimization pipeline it corrupts the optimization itself.

  • Do not report AV-Align alone.

    It is the weakest standalone tracker and the most preprocessing-sensitive metric in the set (CV 0.54, temporal rank-flip 0.66, PEAVS τ 0.00). If it appears at all, it should appear beside metrics that behave differently.

  • For temporal synchronization, prefer Synchformer/DeSync — and say that is what you measured.

    It is far and away the best offset tracker (τ = 0.84, temporal rank-flip 0.19) and the only metric that separates close generator checkpoints. But its competence is axis-specific: it is weak on content-disruption families and agrees with the perceptual proxy at only τ = 0.07.

  • Do not expect a meta-metric to paper over the gap.

    Neither ridge fusion nor a k-NN combiner beats the best single metric against PEAVS. There is no free combination of these four scores that recovers a human-aligned one.

  • Report a Reliability Card, not one number.

    Metric-family scores, uncertainty, and rank resolution — the seven axes above. And check whether the gap you are claiming exceeds the metric's minimum detectable difference before you claim it.

Limitations

What this audit does not settle

Scope of the clip pool. The primary audit is 75 AVSync15 clips, a curated, highly-synced pool. A 150-clip replication and a 75-clip in-the-wild VGGSound rerun both reproduce the ordering, and absolute τ is lower in the wild as expected — but this is not a claim about every AV domain.

PEAVS is a proxy, not human labels. Perceptual grounding uses the PEAVS metric (human-aligned, Pearson 0.79) as a scalable stand-in for fresh human ratings. PEAVS is itself a learned model and may share representational biases with the embedding metrics, so this axis is read as PEAVS agreement, not perceptual ground truth. Direct human preference annotation remains future work before any claim of perceptual superiority.

Controlled distortions, not generative artifacts. The oracle imposes known synthetic desync, which is what makes it annotation-free. The MMAudio study complements it with real generated outputs, but it is a leaderboard-resolution stress test on 45 clips, not a full generator benchmark.

Fusion was tested, not exhausted. Ridge and k-NN combiners fail here; supervised calibration against direct human labels is untested and remains open.

Citation

BibTeX

Peer-reviewed and accepted as a poster; cite it as a workshop presentation. Authors retain copyright; the paper is released under CC BY 4.0.

cite this work
@misc{sharma2026avsync,
  title        = {What Do Audio-Visual Synchronization Metrics Actually Measure?},
  author       = {Sharma, Jai Kumar and Tapadiya, Peeyush},
  year         = {2026},
  howpublished = {Gen4AVC Workshop at ECCV 2026},
  note         = {Workshop poster}
}