A Coding Guide to Google Research’s MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
In this tutorial, we work with MSEB, the Massive Sound Embedding Benchmark from Google Research, and approach it from the perspective of what a leaderboard number actually means: the evaluator surface. We install the package and map its three layers, then write two deliberately different encoders against the frameworks own abstract base class: one that measures loudness over time and one that measures timbre, and encode a small synthetic corpus we generate in the notebook so nothing has to be downloaded. We drive the classification, clustering, retrieval, and segmentation evaluators over those embeddings, call the metric functions directly to see what each one rewards, and finish by assembling the TaskMetadata a real submission carries. The result is a comparison in which the two encoders...
A Coding Guide to Google Researchs MSEB: Writing Sound Encoders to the Benchmark Contract and Scoring Them Across Classification, Clustering, Retrieval and Segmentation
In this tutorial, we work with MSEB, the Massive Sound Embedding Benchmark from Google Research, and approach it from the perspective of what a leaderboard number actually means: the evaluator surface. We install the package and map its three layers, then write two deliberately different encoders against the frameworks own abstract base class: one that measures loudness over time and one that measures timbre, and encode a small synthetic corpus we generate in the notebook so nothing has to be downloaded. We drive the classification, clustering, retrieval, and segmentation evaluators over those embeddings, call the metric functions directly to see what each one rewards, and finish by assembling the TaskMetadata a real submission carries. The result is a comparison in which the two encoders trade places depending on which evaluator is asked, which is the argument for a multi-task benchmark made in numbers rather than in prose.
Copy CodeCopiedUse a different Browser
We install mseb and import the three layers that a benchmark run walks down. The types module holds the shapes every task speaks, Sound, SoundEmbedding, Score and TaskMetadata; the encoder module holds MultiModalEncoder, the contract our own model implements; and the evaluators package holds one module per task family. We import only the four evaluators this notebook drives, because the classification, clustering, retrieval, and segmentation modules depend on nothing heavier than NumPy and scikit-learn. In contrast, the reranking and transcription evaluators pull in Whisper and the task runner pulls in TensorFlow and apache-beam. Everything below therefore runs on a free CPU runtime with no dataset download and no accelerator.
We start with the type contract, because every other layer is expressed in it. A Sound carries a waveform, along with SoundContextParams, the identifier, sample rate, length, language, and optional transcript, which follow the audio through the whole pipeline. A SoundEmbedding carries an array of N embeddings and an array of M timestamp pairs, and the relation between N and M is the benchmarks vocabulary: M equal to N means one vector per frame, while M equal to one means a single utterance-level vector, which is what our encoders produce. EncodingStats records the input and embedding sizes and exposes compression_ratio, here a thousandfold reduction from audio to vector. A Score is a metric name, a value and its bounds, and it validates itself at construction, rejecting an empty metric name or a minimum above its maximum, so a malformed number cannot reach a leaderboard. The embedding field also accepts N strings instead of N vectors, which is the door that step 8 walks through.
We write two encoders by subclassing MultiModalEncoder, whose abstract methods are exactly three: _setup loads whatever the model needs, _check_input_types rejects anything that is not a Sound, and _encode turns a batch into SoundEmbedding objects. The framework owns setup and encode, and encode is what attaches EncodingStats to every result, so our code never fills that in by hand. EnergyEnvelopeEncoder averages energy in sixteen equal time slices and therefore describes only how loudness moves; SpectralProfileEncoder pools the mean log-magnitude spectrum into sixteen bands and therefore describes timbre. Both L2-normalise their output so a dot product is a cosine. Encoding one decaying note through each shows the difference immediately: the envelope encoder sees the decay, and the spectral encoder sees a single peak at 440 Hz.
We synthesize a corpus in which two cues are deliberately separated. The spectrum says which class a clip belongs to, a tone, a chirp or noise, while the amplitude envelope is drawn per item and is independent of class, so it identifies which clip it is without saying anything about what it is. We normalize every waveform to unit RMS before applying the envelope, leaving the envelope as the only loudness cue. We render each of the thirty-six items twice, once as the document and once as a noisier second take of the same clip, and encode both sets with both encoders into MSEB embedding caches, the plain dictionaries from sound id to SoundEmbedding that every evaluator consumes. The mean same-class and other-class cosine similarities printed here read as a prediction about the next three steps: only the spectral encoder separates the classes at all.
ClassificationEvaluator takes a table of class embeddings as its weights and a distance function, and we build the weights as class prototypes, the mean unit vector of each class. Its two methods separate cleanly: compute_predictions returns a raw score per class for every cached embedding, and compute_metrics turns those together with ClassificationReference labels into the list of Score objects that a leaderboard stores. Setting top_k_value to two adds Top-2 Accuracy alongside accuracy, balanced accuracy and the weighted precision, recall and F1. The spectral encoder classifies the corpus perfectly, and the envelope encoder lands well above chance but far below it, which is the ordering the cosine gap predicted.
ClusteringEvaluator asks the harder version of the same question, because it never sees a label at encode time: it runs KMeans over the cache. It scores the clusters against the labels with V-measure, the harmonic mean of homogeneity and completeness. The gap between the two encoders widens sharply here compared with classification, because a supervised prototype readout can exploit a faint cue that unsupervised clustering cannot find on its own. One practical detail is worth copying into any reproducible benchmark run: the evaluator constructs MiniBatchKMeans without a random_state, so it falls back to NumPys global generator, and without seeding that generator an unstructured embedding space scores anywhere between roughly 0.01 and 0.08 from run to run.
RetrievalEvaluator answers a different question from the two before it, and we set the task up so that difference is visible. Each query is the noisier second take of exactly one document, so the target is identity rather than category. We index the document embeddings in a BruteForceSearcher, ask for predictions over the query cache, and pass one RetrievalReferenceId per query naming its single correct document. The evaluator returns MRR, exact match, recall at our top_k and NDCG at ten. The result inverts the previous two steps: the envelope encoder retrieves every clip at rank one, because the envelope is an item fingerprint. In contrast, the spectral encoder ranks slightly worse because clips of the same class look alike to it. The printed top-five lists make the mechanism plain, one neighbourhood class-random and the other class-pure.
We call the metric functions directly, without an evaluator around them, because they are the layer the task families share. compute_word_errors and compute_character_errors take two strings and return errors and totals separately, so the caller decides how to aggregate a corpus. The ranking metrics take a reference and a ranked list of identifiers, and comparing exact match, reciprocal rank and nDCG over the same ranking shows what each one pays for position. One sharp edge is worth naming: compute_ndcg_at_k assumes a single relevant document and compares it by equality, so passing a list of relevant ids silently scores zero. In contrast, MRR, which does accept a list, still looks correct. We close with compute_lp_norm and compute_dynamic_time_warping_distance, the embedding-space distances behind the reconstruction and stability tasks.
SegmentationEvaluator scores what was said and where it was said as separate quantities, and it uses the string form of SoundEmbedding that step 1 mentioned: the embedding array holds one term per segment and the time