Knowledge work with recordings

Video Metadata and Tagging for a Library You Can Search

Video metadata is the information about a video that the picture and sound do not carry: what it is about, who appears in it, how it was encoded, and who owns the rights. A useful library records three kinds (descriptive, technical and rights), stores them where they will survive copying and export, and tags videos from a fixed list of terms. Transcripts make tagging faster because they show what is actually discussed, minute by minute.

9 min read · Updated

The three kinds of video metadata

Metadata work goes wrong when every field gets lumped into one list. It helps to separate fields by the question they answer:

Descriptive
What is this video about? Title, summary, people, places, topics, event, date filmed, spoken language, audience.
Technical
What is this file? Container, video and audio codecs, resolution, frame rate, duration, audio tracks, file size.
Rights
What may we do with it? Owner, licence, release forms, music and stock footage licences, restrictions, expiry dates, credit line.

Descriptive metadata is what people search by, and it takes human judgement. Technical metadata can be read automatically from the file with tools such as MediaInfo or ffprobe, so there is little reason to type it by hand. Rights metadata is the field set most often skipped and most often regretted: a beautiful interview is useless for a campaign if nobody knows whether the subject signed a release or whether the music licence covers paid advertising.

Choosing descriptive fields that earn their keep

Every field you add is a field someone has to fill in for every video, forever. Start with the questions people actually ask when they look for footage: "do we have anything with the CEO talking about sustainability?", "what did we film at the 2025 conference?", "is there a Spanish-language version?". Fields that answer those questions stay; fields nobody searches go.

A compact starting set for most organizations:

  • Title, written for humans rather than copied from the file name
  • Summary of two or three sentences
  • People appearing or speaking, using one agreed form of each name
  • Topics, chosen from a controlled list
  • Event or project, and date filmed
  • Location
  • Spoken language, plus the languages of any subtitle files or translated versions
  • Content type: interview, presentation, B-roll, tutorial, advert

If you need a reference point, general-purpose schemes exist. Dublin Core defines a small set of broad elements such as title, creator, subject, description, date and rights, and is widely used in libraries and repositories. Media-specific schemes, such as PBCore in public broadcasting or the IPTC Video Metadata Hub in news, go into much more detail. Borrowing field names from a scheme makes it easier to move records into another system later, even if you never adopt the whole standard.

Embedded versus sidecar metadata

There are two places to keep metadata, and each fails in a different way.

Embedded metadata lives inside the video file. MP4 and MOV files have slots for a title, comments, creation date and more; MKV has its own tag system; editing applications can write XMP data into many formats. The advantage is that it travels with the file. The drawbacks are that support varies between applications, many exports and conversions drop or rewrite it, and many video platforms strip or ignore most of it on upload. Embedded fields are also awkward to search across a whole library.

Sidecar metadata lives outside the file: an XMP sidecar file next to the video, a spreadsheet catalog, or the database of a digital asset management system. It is easy to search, edit in bulk and back up, and it survives any re-export of the video. The risk is separation: rename or move the video without its sidecar and the link breaks.

A practical rule for most libraries is to treat the catalog as the source of truth and embed only a few basic fields, such as title and copyright notice, as a fallback. Give each video a stable identifier, use it in the file name and in the catalog row, and never reuse it. The same logic applies to audio collections; see organizing audio files for naming and catalog patterns that carry over directly.

Controlled vocabularies: one word for one thing

Free tagging feels fast and produces a mess. Within a year the library has "customer story", "case study", "testimonial" and "client video" for the same kind of content, plus "NYC", "New York" and "new-york" for one city. Searches miss half the results and nobody knows which tag to use.

A controlled vocabulary is simply an agreed list of terms, with one preferred term for each concept. To build one:

  1. Pull the tags people already use, if any, and group the synonyms.
  2. Pick one preferred term for each group, and record the others as "use instead" pointers.
  3. Keep the list short at first. A few dozen topic terms cover a surprising amount of a typical corporate or nonprofit library.
  4. Organize terms in a shallow hierarchy where it helps, such as Events, then Conferences, then a specific conference.
  5. Name one owner who approves new terms, and review the list once or twice a year.
  6. Apply the same rule to people's names: one spelling, one order, one form for every appearance.

Large public vocabularies, such as library subject headings, show how far this can go. Most organizations need a small fraction of that, kept current.

Drawing tags from transcripts

The slowest part of tagging is knowing what is in the video. Watching an hour-long recording to tag it is not realistic for a large backlog. A transcript lets you read or skim the content in a few minutes and see where each topic comes up.

A workable routine:

  • Skim the plain transcript once and note the main subjects in your own words.
  • Map those subjects to terms from your controlled vocabulary rather than inventing new tags.
  • Search the transcript for names of people, products and places, and add each one in its agreed form.
  • Use the timestamped transcript to note where major segments start, so a catalog record can say "pricing discussion at 18:40" rather than just "pricing".
  • Check every name against the recording before it becomes a tag. Speech recognition often misspells names, and a wrong name in a tag spreads to every search.

Word counts can suggest candidates, since a term mentioned forty times is probably a topic, but the final tags should be chosen by a person. The most frequent words in a transcript are rarely the most useful tags. Segment-level timestamps also make later reuse easier; the clip-hunting approach in finding clips in long videos builds on exactly this.

Tagging a nonprofit's back catalog

A hypothetical charity has around 600 videos from ten years of events, appeals and field visits, tagged inconsistently or not at all. A two-person team agrees a list of 45 topic terms, a name list for staff and partners, and six content types. They transcribe the videos people request most often, skim each transcript, and fill in a spreadsheet row with summary, people, topics, language and rights status. Videos with no rights information are flagged "do not reuse" until someone confirms a release exists. After three months, about a third of the library is catalogued, and that third covers most of what colleagues actually ask for.

Rights metadata deserves its own process

Rights information rarely lives in the video and often lives in someone's inbox. Collect it at the time of filming, not years later. For each video, record who owns it, whether the people who appear signed releases and what the releases cover, the licences for music, stock footage and fonts, any territory or time limits, and the date a licence expires.

Treat "unknown" as a value. A video marked "release status unknown" is honest and searchable; a blank field looks like permission. Licence terms vary widely, so read the actual agreement before reusing material, and ask your legal team when in doubt.

Pitfalls of video metadata projects

  • Designing fifty fields and filling in six. Start small and add fields when someone needs them.
  • Copying the file name into the title field. "IMG_4471.MOV" tells a searcher nothing.
  • Relying on embedded tags that a single export can wipe.
  • Letting anyone add new tags without review, which recreates the synonym problem within months.
  • Tagging every video exhaustively before anyone uses the library. Prioritize the material people request.
  • Ignoring language. If translated versions exist, record which language each file is in and link the versions to one another. Translating platform titles and descriptions is a separate task, covered in translating YouTube titles and descriptions.

Where mydubly fits in a tagging workflow

mydubly produces the transcripts that tagging draws on. You choose a video from your device in MP4, MOV, WebM, MKV or M4V format, up to 2 hours long. The audio is extracted in your browser and sent as compressed audio chunks for recognition, while the picture stays on your device. You download a plain transcript.txt for skimming, a timestamped transcript with [m:ss] labels for noting segments, and SRT and VTT subtitle files if the video needs captions. The spoken language is detected automatically, and you can request a translated transcript in one of 21 languages in the same job. See video to text for the details, or the MOV to text page for camera and phone footage.

mydubly does not read, write or edit metadata, suggest tags, identify people or analyze the picture. It does not store or catalog your library: uploaded audio and results are deleted within 30 minutes of a job finishing, so save the transcripts with your catalog. If you create a translated or dubbed version, treat it as a new asset with its own record and check what metadata the new file carries rather than assuming the original's tags came along.

Transcripts cost 1 credit per minute, minimum 5 credits per file. Transcribing a 30-minute event recording is 30 credits (3¢).

Next step: a pilot of twenty videos

Pick twenty videos people ask for often. Draft the field list and a first controlled vocabulary, transcribe the twenty, and catalog them. Then ask two colleagues to find specific footage using only the catalog. Wherever they struggle, adjust a field or a term before you scale up. The broader case for searchable transcript archives, written for course libraries, is in making lecture recordings searchable.

Frequently asked questions

What is the difference between video metadata and tags?

Tags are one kind of descriptive metadata: short terms that say what a video is about. Metadata also covers longer descriptions, technical details such as codec and duration, and rights information such as licences and releases. A good library uses tags from a controlled list alongside the other fields.

Do video platforms keep the metadata embedded in my files?

Often not. Many platforms strip or ignore most embedded metadata when you upload, and use the title, description and tags you enter in their own interface instead. Keep your full record in your own catalog and check each platform's current documentation for what it reads.

How many tags should each video have?

Enough to answer the searches people actually make, and no more. For many libraries that is a handful of topic terms, the people who appear, the event or project and the content type. Long lists of loosely related tags make search results noisier, not better.

Can metadata be generated automatically?

Technical metadata can, because tools read it straight from the file. Descriptive metadata can be drafted with help from transcripts and other automated tools, but a person should choose the final tags and check names. Rights metadata depends on contracts and releases, so it always needs a human.

Should I use an existing metadata standard?

Borrowing field names from a standard such as Dublin Core is a sensible default because it eases moving records between systems later. Full media-specific standards are worth adopting if you exchange records with broadcasters, archives or news agencies that use them. Small teams usually do fine with a short custom list built on standard names.

How do I tag videos that are in different languages?

Record the spoken language as a field on every video, list the languages of available subtitle files, and link translated versions to the original's record. Keep topic tags in one working language so a single search finds all versions. A translated transcript can help a cataloguer who does not speak the video's language understand what it covers.