# Interview Transcription in DaVinci Resolve – Workflow for Claude

As of: October 2026 · tested with DaVinci Resolve Studio 21 · used on the documentary "Such(t)en" (15 interviews, Swiss German) and on an English test interview
By Kamil Goerlich, film editor · https://videograf-bodensee.de

**How to use this file:** attach it in a conversation with Claude and write: "Please follow this workflow for my interviews."
"User" below means you, the person giving Claude the task.

**Requirements**
- **DaVinci Resolve Studio** (the free version does not allow external scripting)
- **Claude in the desktop app** with access to your computer and **screen control** (for the subtitle menu)
- A **connection between Claude and Resolve** via the Resolve scripting interface (MCP server), so Claude can run scripts in Resolve
- Python with ReportLab for the PDFs (Claude can set this up in its own workspace)

Menus and script commands may change with new Resolve versions. Use at your own risk; always make a project backup first.

This file describes how Claude transcribes interviews in a Resolve project, prepares the audio
for it, creates PDFs, names the timelines and builds best-of timelines.
**Ground rule: originals are never changed.** Claude only works with copies.
This is about editing and organisation in Resolve, not about publication: for the editor **everything** is relevant.

---

## 0. Questions before starting (always ask, collected in one message, as multiple choice)

Ask every question with numbered answers so the user only has to reply with numbers (e.g. "1, 1, 2, 1").
If a tool for choice questions is available (click answers), use it.

1. **What should be processed?**
   1) the timeline currently open
   2) all timelines in a bin
   - On **1** continue directly. The target bin is then `<bin of the timeline>_TRANSCRIPT`.
   - On **2** ask for the bin name, and whether all timelines in it should be processed
     (1 yes, 2 no → which ones to skip).
2. **Which language is spoken?** (for speech recognition)
   1) English 2) German 3) French 4) Italian 5) Other
   - On **5** ask which one (for a dialect such as Swiss German: "German" + Extended Language Support).
3. **Which language should subtitles, PDF and quotes be in?**
   1) English 2) German 3) French 4) Italian 5) Other
   - On **5** ask which one. If it differs from question 2, the **subtitles are translated** (step 4b).
4. **Remove background noise (noise reduction / Voice Isolation)?**
   1) Yes 2) No
   - Yes → Voice Isolation **60**, with strong background noise (outdoors, street, restaurant) **80**. No → off. See step 3.5.
5. **Where should the PDFs go?**
   1) Suggestion: `<project folder>/Transcripts`
   2) another folder → ask

Question 2 must be answered. If other answers are missing: question 1 → 1, question 3 → same language as question 2,
question 4 → Yes, question 5 → 1. If the user does not answer at all, ask again instead of guessing.

**Do not ask — decide or find out yourself:**
- **Target bin:** always `<source bin>_TRANSCRIPT`, on the same level as the source bin.
- **Names of interviewer and interviewee:** determine them from the subtitles (greeting, form of address, introduction).
  Not found → use the role ("Interviewer", "Photographer" …) and ask the user at the end.
- **Audio specifics:** analyse automatically (see step 3.1), e.g. whether both stereo channels are identical.
- **Best-of timelines:** always built (bin `ZITATE`).

---

## 1. Preparation

- Check that Resolve is open and the correct project is loaded (`project.GetName()`).
- Recommend to the user to export a project backup beforehand (Project Manager → right-click → Export, .drp).
- List all timelines in the source bin, with length, number of video/audio tracks, channel layout and clip names.
  Briefly show the user the list (name, length, detected cameras/mics) before processing begins.
- Order: short interviews first; the longest last. Start with **one** interview completely
  and show the user the PDF before the rest follow (have the format approved).

---

## 2. Copy timeline

- `tl.DuplicateTimeline("<original name>_TRANSCRIPT")`, move the copy into the target bin `<source bin>_TRANSCRIPT` (`mp.MoveClips`).
- The original remains untouched, the copy is 1:1 timecode-identical (important for the later edit).

---

## 3. Prepare audio (only in the copy, only for the transcription)

Goal: clean, clearly intelligible, **equally loud mono tracks**, so that speech recognition works as accurately as possible.

1. **Analyse the audio automatically** (instead of asking the user): for each audio track read the source file
   (`run_script_unsafe`, Python `wave`) and measure the RMS level per channel and per second, plus the level of the difference L−R.
   - Difference very quiet (> 30 dB below the signal) → **channel 1 and 2 identical** (same mic).
   - Channels alternately loud/quiet → **two different mics** (e.g. interviewee left, interviewer right).
     Who speaks on which channel follows from content and level curve (questions = interviewer).
   - Several cameras: take the track with the best signal-to-noise ratio (default: top audio track A1).
   - Briefly state the result in the interview list from step 1.
2. **Convert stereo tracks to mono:** lay the clip again onto separate mono tracks
   (`AddTrack("audio","mono")`, `AppendToTimeline` with `mediaType: 2`, then `SetSourceAudioChannelMapping` → `"type":"mono"`, `channel_idx` per channel).
   - `endFrame` = the original's source end frame (not −1), otherwise every segment is 1 frame short. Compare start/end with the original afterwards.
   - **Channel 1 and 2 identical** (common with camera audio: one stereo track, left and right the same): make exactly **one** mono track from it (channel 1), do not use the second channel.
   - **Two different mics:** **two separate mono tracks**, do not mix, name the tracks after the speaker.
   - **Same mic, levelled differently:** take the clean channel (e.g. the quieter one in case of clipping).
3. **Delete old audio tracks** (`DeleteTrack("audio", 1)`), so that the prepared tracks are **at the very top**.
   ⚠ Subtitle recognition listens to the top audio tracks, **even if they are disabled/muted**.
   If old tracks stay on top, only nonsense comes out (e.g. hundreds of times "Vielen Dank.").
4. **Match loudness – every track, every channel separately:** normalize all segments
   (`NormalizeAudioLevel`, mode ITU-R BS.1770-4, **−20 LKFS**, `NORMALIZE_AUDIO_SET_LEVEL_INDEPENDENT`).
   ⚠ **Check afterwards** that the level really changed (`GetProperty("Volume")` of the clips).
   In the test run the volume stayed at 0 dB: the **quiet channel** (interviewer, approx. 20 dB quieter) was **not raised**.
   Reason: normalization measures the **whole stereo file**, not the mapped channel.
   So take the difference from the measurement in 3.1 and raise the quiet channel with
   `item.SetProperty("AudioVolume", <value of the loud channel> + <difference>)` (tested, works).
   Goal: both speakers about equally loud. In the test run this also clearly improved recognition (gap closed).
   Quiet recordings sometimes need +9 to +30 dB, that is normal.
5. **Voice Isolation** according to the answer to question 4 (`SetVoiceIsolationState(track, {"isEnabled": True, "amount": 60})`):
   Yes → **60**, with strong background noise (outdoors, street, airport, restaurant) **80**; No → leave it off.
   The amount may be high, because this audio only serves the transcription and the best-ofs; the original stays untouched.
6. Long scripts can run into a timeout: **before a repeat, always first check how far it got.**

---

## 4. Create subtitles

### 4a. Recognition (in the spoken language)
- In the `_TRANSCRIPT` timeline via the menu **Timeline → AI Tools → Create Subtitles from Audio…** (via screen control).
- Tick **"Enable Extended Language Support"**, then select the **language** (from 2a)
  (it jumps back to "Auto" after ticking). Check via zoom before "Create".
  Without Extended, Swiss German is unusable.
- The language list can only be operated with **full-screen control** (blocked in the background):
  open the list, type the language name (jumps to it), click it.
- ⚠ Resolve sometimes starts recognition **immediately without a dialog**, using the last settings.
  Then check via `project.GetCreateSubtitlesFromAudioStatus()` and verify the language in the result.
- Wait for the result (query `GetItemListInTrack("subtitle", 1)` until entries are there).
- **Plausibility check:** number of subtitles, most frequent texts. Many identical phrases = recognition heard silence → check step 3.3.
- Mark gaps in the text (sentence breaks off, next one starts mid-sentence) with **[ ]** and report them to the user.

### 4b. Translation (only if 2b ≠ 2a)
Resolve does not translate by itself, and the subtitle text **cannot** be changed via script (`SetName` on subtitles fails).
Therefore:
1. Translate the cleaned-up subtitles into the output language, **same timecodes**, and save as SRT
   (`<timeline name>_<language>.srt`, e.g. `…_DE.srt`, in the PDF folder).
2. Place the SRT as a **second subtitle track** at the start of the timeline (tested, frame-accurate):
   - via script: `tl.AddTrack("subtitle")`, import the SRT into the target bin with `mp.ImportMedia([path])`
     (write SRT times relative from 00:00:00,000, not from 01:00:00);
   - via screen control (full screen): drag the SRT from the Media Pool onto the new subtitle track at the **start of the timeline**;
   - then compare count and start frame with the original track and **rename** the track (the name resets on drag).
   - The menu File → Import → Subtitle opens a system file dialog that cannot be controlled → do not use.
   ⚠ **Not** via `AppendToTimeline` into an existing timeline: that ignores `recordFrame` and `trackIndex`
   and appends the subtitles to the **end of track 1**.
3. The original track (spoken language) is kept; name the tracks (e.g. "EN original", "DE").

---

## 5. Read out and clean up text

- Read out all subtitles with `GetStart()`, `GetEnd()`, `GetName()` and save as SRT/text.
  Calculate timecodes from timeline start frame and frame rate (do not hard-code).
- Clean up in the output language (question 2b): correct recognition errors, names, places, technical terms.
- **Assign speakers** (Resolve does not recognize speakers): based on the content and, with separate mics, on the level per channel.
- Determine **names** from the text (see section 0).
- Mark uncertain passages with **[ ]**. Collect uncertain names and list them for the user to check at the end.

---

## 6. PDF per interviewee

Created with Python/ReportLab (`build_pdf.py` + one `<name>_content.py` per interview). Structure:

1. **Header table:** person, shooting date, location, interviewer, length, timeline name, note on markings
2. **Heading:** a short, fitting title for the interview (also used for the timeline name, see 7)
3. **Summary** (a few paragraphs)
4. **Strongest quotes** (approx. 10–15): timecode, quote in «…», topic keyword
5. **Topics:** table topic / content / passages (timecodes)
6. **Cleaned transcript** with timecodes and speakers

- Timecodes = **timeline timecodes** of the `_TRANSCRIPT` timeline (identical to the original).
- Before storing, check that every quote timecode occurs in the transcript.
- File name: `YYMMDD_Interview-<Name>.pdf`. Storage: folder from question 3.

**Transcribe completely:** everything is relevant for the edit, **nothing is left out** –
including breaks/preliminary talks, private passages and statements about third parties.
Only mark as a **note for the editor**, do not shorten:
- **"SENSIBEL"** (sensitive) for suicide, violence, abuse, third-party addiction etc.
- **"PAUSE / VORGESPRÄCH"** (break / preliminary talk) for conversations outside the actual interview.
- **"NICHT VERWENDEN (Bitte im Interview)"** (do not use – asked in the interview) when someone explicitly asks for it in the interview.

---

## 7. Rename transcript timelines

After completing the PDF, rename the `_TRANSCRIPT` timeline:

`YYMMDD_<heading from the PDF>_TRANSCRIPT`

- **YYMMDD = shooting date from the camera metadata** of the clip with the good audio
  (`GetClipProperty()` → "Date Recorded", otherwise "Shot Date" / "Date Created").
  If the audio file only carries a file date (e.g. date of copying), take the date of the picture clips;
  if in doubt, the file name or ask the user.
- Keep the heading short, replace spaces with hyphens, no special characters except hyphens, write out umlauts.
- The original keeps its name.

---

## 8. Best-of timelines (bin `ZITATE`, always)

- Create bin `ZITATE` (if not present).
- One timeline per interviewee `<firstname>-<lastname>-best-of` (e.g. `jeffrey-spoeri-best-of`), lowercase, umlauts written out.
- Content: the strongest quotes from the PDF in PDF order, each with about 1 second of room before and after.
  If the ranges overlap, merge them into one clip. Do not cut into the next question.
- **Picture:** the original camera clips. **Audio:** the **prepared audio** from the `_TRANSCRIPT` timeline
  (mono tracks per speaker, raised quiet channel, normalization, Voice Isolation) – **not** the original audio.
  Calculate source frames via the mapping timeline timecode → source timecode (picture from the original, audio from the `_TRANSCRIPT` timeline),
  split ranges at edit points.
  ⚠ Newly inserted clips do not take over the settings automatically. Set again:
  channel mapping (`SetSourceAudioChannelMapping(old.GetSourceAudioChannelMapping())`), volume (`SetProperty("AudioVolume", …)`) –
  both can be copied from the transcript clip – and Voice Isolation per track **explicitly** with `{"isEnabled": True, "amount": 60}`
  (copying the state from the transcript timeline fails silently). Then check against the `_TRANSCRIPT` timeline.
  If there is no interview picture (e.g. only B-roll with speed changes), take the audio only and tell the user.
- Set a **marker** with the quote text on each quote (SENSIBEL quotes with a red marker).
- **Take the subtitles along** (in the output language): create an SRT with the new positions and import it.
  Source: the output-language subtitles from the `_TRANSCRIPT` timeline.
  Working order: **create empty timeline → add subtitle track → insert the SRT first via `AppendToTimeline`**
  (it then lands at the start), **then** insert audio/picture with `recordFrame` and set the markers.

---

## 9. Final message to the user

Short and clear:
- Which interviews are done, where the timelines (bins), PDFs and SRT files are located
- **Please check:** determined names, uncertain passages [ ], marked sensitive passages
- What was deliberately not done (e.g. skipped interviews, best-of without picture)

---

## Lessons (short version)

- Subtitle recognition listens to the top tracks, even disabled ones → prepared tracks on top, delete old ones.
- Extended Language Support is mandatory for dialect; (re)select the language after ticking; language list only works in full screen.
- Channel 1 and 2 are often identical → then only one mono track. Separate mics → two tracks **and** raise the quiet channel.
- `NormalizeAudioLevel` measures the whole stereo file, not the mono channel → raise the quiet channel via `AudioVolume`, then check.
- Resolve does not translate subtitles; subtitle text cannot be changed via script → import the translated SRT and drag it to the start as a second track.
- `AppendToTimeline` with an SRT appends at the end of the timeline; only in an empty timeline does it land at the start.
- Transcription and best-of use the prepared audio (may be filtered heavily); the original stays untouched.
- Transcript timelines are 1:1 copies → timecodes in PDF, transcript and original match.
  Therefore do not move or trim the originals any more.
