Project DoLittle | Data update
A Careful First Pass at an Orca Acoustic Corpus
What it takes to turn public hydrophone archives into an auditable research asset: source-aware retrieval, exact 30-second locality, quality triage, and a conservative listening-cleanup path.

A usable marine-audio dataset is more than a pile of recordings. It needs to preserve where each sound came from, how it relates to neighboring audio, what its licensing terms permit, and how much uncertainty remains before it is used in a model. Our Orca work is building that foundation first.
What We Have So Far
We have completed label-guided or annotation-led Orca context acquisition across DORI, SanctSound, OOI, and several DCLDE collections. Each retained item is an exact 30-second raw context with its parent recording order, time offset, source object or URL, and license partition. Multi-context passes also preserve neighbor relationships so future work can reconstruct chronology rather than treating every clip as an isolated sound.

DORI-ONC is currently the largest partition, followed by DCLDE DFO-CRP, JASCO/VFPA, DFO-WDLP, and DORI-Orcasound. The differences are informative, but they are not a leaderboard for where Orcas are most common: coverage, annotation practice, recording length, and public access all shape these counts.
Why 30 seconds? It is long enough to preserve local acoustic context and neighboring-call structure, but small enough to retain, audit, and later process reproducibly without downloading full multi-gigabyte archive recordings.
Discovery Is Not Acceptance
Several additional continuous-audio sources have been scanned in bounded pilots. These pilots are deliberately separate from the retained Orca corpus. A detector-positive score window is a lead for review, not a biological conclusion or a dataset row.

This distinction matters. MBARI Pacific Sound and Ogasawara pilots, for example, generated many detector score windows, but those windows remain pending source-specific review and threshold calibration. Likewise, inaccessible archive objects are recorded as access outcomes, not as evidence that an animal was absent.
Source-Aware Acoustic Quality Assessment
Hydrophone data is wonderfully varied: quiet water, distant calls, steady machinery, handling noise, tonal artifacts, and shifting ambient conditions can all occupy the same 30 seconds. We built an automated evaluator to measure raw-audio noise burden using temporal variation, band energy, spectral texture, low-frequency hum, persistent tones, impulsive changes, and frame-level noise-floor features.
The evaluator estimates a source-relative risk of moderate_or_worse noise. It does not identify species, prove a vocalization is present, or automatically decide whether a context is suitable for tokenization. Human listening still confirms the meaningful tiers: quiet/soft, moderate, or loud/repetitive noise.

A Conservative Cleanup Path for Listening
For confirmed quiet/soft candidates, and in a separate experimental partition for confirmed moderate-noise candidates, we now have a working listening-cleanup derivative. It uses the released Earth Species Project Biodenoising DNS48 checkpoint at its native 16 kHz output, applies a narrow 4 kHz Q35 notch to remove a persistent model residual, and peak-normalizes the result to 0.90.

We compared stationary spectral subtraction, a source-reference bounded spectral-gain method, DNS48 with different normalization strategies, an optional 4 kHz notch, and hybrid outputs that restored the original recording above 8 kHz. Listening showed that DNS48 could remove substantial noise while leaving audible whale vocalizations. The 4 kHz notch became necessary after peak normalization made a pre-existing narrow model residual noticeable.
The hybrid high-band experiment was useful precisely because it did not deliver a dramatic result. Across 16 paired clips, the restored band above 8.25 kHz was a median 1.10% of total low-plus-high energy and about 19.5 dB below the cleaned low band. We cannot yet tell whether that residual high-band structure is biological, environmental, or instrumental. The cleaner model-only notched output is therefore the working default; the hybrid remains an optional archival-fidelity derivative.
Scope boundary: DNS48 changes the acoustic representation substantially and emits 16 kHz audio. It is not a raw replacement, it is not valid input to our legacy 44.1 kHz Orca detector, and it is not yet a tokenization authorization. Its current role is to create an auditable, research-only listening and preparation derivative.
From Source Recording to a Prepared Derivative

- Retain exact local context. Source provenance, license, parent recording order, and 30-second offsets travel with the audio.
- Measure noise burden. The evaluator scores raw-audio features and produces a source-relative review priority.
- Confirm the tier. Human listening distinguishes quiet/soft, moderate, loud/repetitive, no-audible-vocalization, and uncertain outcomes.
- Create a separate derivative. Quiet/soft clips move to the working DNS48 cleanup path; moderate clips stay in a separate cleanup-evaluation partition; loud or uncertain clips remain raw-only holds.
- Freeze an explicit future selection manifest. Any tokenization experiment must retain raw and derived paths, source/license partition, quality evidence, and the decision rule that selected it.
What Is Next
| Next step | Why it matters |
|---|---|
| Freeze a fresh source-stratified quality holdout | Validate a combined evaluator without tuning on the final test set. |
| Expand review in the largest retained partitions | Convert source-relative risk scores into confirmed quality tiers at scale. |
| Validate a detector for the 16 kHz cleanup representation | Our legacy Orca detector expects features through 20 kHz and cannot evaluate DNS48 output fairly. |
| Keep sweeping accessible, well-mapped sources | Orcasound archive access, full SanctSound terms, and OOI historical mapping remain real acquisition constraints. |
| Run a manifest-driven tokenization evaluation | Start with confirmed quiet/soft contexts, keep moderate cleanup candidates separate, and preserve every decision. |
Why Provenance and Uncertainty Matter
The primary outcome so far is an auditable workflow that preserves the route back to the original recording and records uncertainty around every derivative. This supports reproducibility, source-term compliance, and later independent review of model and data-selection decisions.
Method notes: counts are reconciled from the Orca candidate registry as of September 9, 2026. Denoising is based on a published research-only CC BY-NC DNS48 checkpoint. Charts summarize retained context and bounded discovery pilots; they do not estimate population abundance, calling rate, or full archive coverage.