OpenSpeaks Tools Update: Subtitler, Tome, and Bento Go Stable — Meet Media Dashboard

Translate this post

Just over two months ago, we introduced three tools built by and for community language documenters: Subtitler, Tome, and Bento. They were alpha/beta then. We want to update that all three are now stable, thanks to our fellow technical contributors. We’re also introducing a fourth tool, the OpenSpeaks Media Dashboard (OMD), to help documenters see how media files are used across Wikimedia projects.

Subtitler, Tome, and Bento: from alpha to stable

OpenSpeaks Subtitler is our flagship tool and a subtitle editor. In addition to captioning/subtitling, it can now fetch audio, video, and subtitle files from Wikimedia Commons, edit and machine-translate subtitles, and upload the subtitles back to Commons. It has gotten simpler: load a file, caption it, translate it, publish it, without leaving the tool. A side-by-side translation panel with machine-translation assistance is in place, too. Our future plan is to add a feature for automatic speech recognition (ASR)-assisted rough drafts, though ASR is still weak or absent for Indigenous and various local languages.

OpenSpeaks Tome now supports both metadata and licensing agreements. The metadata section has five tabs: Basic Info, Production, People, and Subtitles. The People tab captures what the Oral Knowledge Framework actually asks for: gender, birth year, Wikidata QIDs, and named roles for interviewers, interviewees, transcribers, and reviewers. Import covers OpenSpeaks Archives XML/JSON, Commons wikitext, ELAR’s Lameta/OPEX, and LAC’s BLAM/CMDI; export covers Commons wikitext, OpenSpeaks JSON, Markdown, and HTML. Within the tool, we have made a major addition: a language database mapping ISO 639 codes to standard and endonym names and language varieties/dialects. This list is prepared by cross-checking against Glottolog and Wikipedia. It’s small, but exactly the kind of infrastructure oral-knowledge work has had to improvise around for years. Now, any user can type the first few characters of a language/language variety/dialect name and choose from a list. We plan to use this list in Subtitler and other OpenSpeaks tools, and others can use it too.

OpenSpeaks Bento was released as a stable version on 1 July 2026. Its three original utilities are intact as tabs: Organise renames and organises files correctly within folders; Duration helps assess the duration of individual audio and video files, as well as the total audio and video duration; and Compress helps compress files for easy sharing with colleagues. The v.1.0 added a fourth tab, Analyse Media, which can pull codec, resolution, loudness, and SHA-256 checksum data for archival-quality checks. Bento also adopted Wikimedia’s Codex design system in this release, fixing accessibility issues. It still remains an offline utility and works after being loaded once.

Introducing OpenSpeaks Media Dashboard

OpenSpeaks Media Dashboard v1.2.0

We’re adding a fourth tool to our tool suite—OpenSpeaks Media Dashboard (OMD).

While our other tools serve the language speaker and Wikimedia communities alike, OpenSpeaks Media Dashboard is primarily for Wikimdian-language archivists. It helps Commons contributors visualise the impact of their contributions by showing statistics.

Input a Wikimedia Commons category containing audio, video and images and run it: it queries Wikimedia’s public APIs in real time, showing which pages across Wikimedia projects use those files, how many people viewed those pages each month, in which language editions, and in which Wikimedia projects, and by which file types. You can see monthly pageview charts, a project breakdown, a file-type view, a table of the most-viewed pages, and export the visualisation as SVG or PNG, or the raw query results as CSV. Most importantly, it can produce an analysis report in wiki code that can be edited to create an on-wiki report. Reports for a particular category can be shared via a link or QR code if one prefers to test themselves. The only downside is it queries live and stores nothing between runs, due to not having a server. As Wikimedia’s pageview processing lags by several weeks, the statistics may be slightly outdated, though that is not an issue introduced by this tool. It faithfully transforms what can be queried in a human-readable manner.

Why did we build this? Impact for oral knowledge work begins with counting the files uploaded. But counting readers helps to understand what that knowledge means to the readers. We hope that OMD can help documenters and project coordinators see how their Commons uploads are used across various Wikimedia projects and language editions. Non-language documentation projects that see image, audio and video contributions can use OMD too. We have used several tools, including GLAMorous, GLAMorgan, and GLAM Stat Tool – Cassandra in the past. But analysing the data took longer, so we started working on OMD to enable quick periodic analysis. The tool has its own built-in help tab, just like Subtitler (to be done soon), Tome, and Bento, which have self-contained user help pages. Check out OMD at https://omd.toolforge.org and let us know whether it works or where it breaks.

Thank you, Indic Wikimedia Hackathon Hyderabad

Indic Wikimedia Hackathon 2026 kindly invited us to mentor community developers Arpitha Bhandary and Govind Lal T. L to improve the OpenSpeaks tool suite. During 25–28 June, we joined 56 other fellow Wikimedians at the Indian Institutes of Information Technology (IIIT) Hyderabad. For three days, Arpitha and Govind fixed countless bugs and built new features into Subtitler and Tome. We were joined by Bharathesha A from the Tulu Wikimedia community, who shared many useful insights drawn from his experience documenting Tulu cultural events, whereas Jnanaranjan Sahu, a long-time Odia Wikimedian and an advisor to OpenSpeaks, provided technical guidance and mentorship.

Arpitha wrote afterwards that the hackathon “was more than just a hackathon… it was an experience I’ll always cherish,” and that what stayed with her most was the mentorship: “not once did anyone make me feel discouraged.” Govind called it “three amazing days” and said contributing code that will benefit the community was a rewarding experience.”

Bharathesha supported the project as an editor throughout, giving Arpitha and Govind the community-side read on what documenters actually need. Jnanaranjan Sahu co-mentored the OpenSpeaks track with me. Though Ranjith Siji, the lead developer for Subtitler, couldn’t attend in person, he reviewed and merged the pull requests that came out of the hackathon. Opino Gomango, our OpenSpeaks Fellow for eastern Indian languages, also shared input to improve Subtitler, as he has been one of the earliest testers of the tool since its inception.

What we need from the Wikimedia community

  • Testers: try Subtitler, Tome, Bento, and Media Dashboard with real recordings and real Commons categories, and tell us what works, what breaks and what can be improved.
  • Wikimedians working with audio, video, or oral history on Commons: tell us where these tools fit into your workflow, and where they don’t and why.
  • Developers: Subtitler and Tome still have open questions around ASR and deeper Commons/Wikidata integration. You can help us integrate them.
  • Community documenters and coordinators: if you’re writing an impact or grant report, try OMD in your project’s Commons category and tell us whether the numbers are actually useful to you.

You can do all of the above on the discussion page of our tool page on Meta-wiki.

Links

Can you help us translate this article?

In order for this article to reach as many people as possible we would like your help. Can you translate this article to get the message out?