Share.

    5 Kommentare

    1. jamiepinheiro on

      I was curious as to how TV/movie character speaking-time was distributed, so I built a tool to ingest datasets of subtitles and calculate the split. I also added „cumulative speaker-share plots“ across time, which help visualize how an episode/film’s speakers change throughout the content. For some titles, speaker data wasn’t available, so LLMs were used to infer the speaker with ~reasonable success (please feel free to flag any mistakes).

      Some interesting findings:

      * Michael, Dwight + Jim make up ~half of [The Office](https://top-talkers.jamiepinheiro.com/title/the-office?shared=true)
      * [Friends](https://top-talkers.jamiepinheiro.com/title/friends?shared=true) has a pretty even split between the leads

      Check it out here! [top-talkers.jamiepinheiro.com](http://top-talkers.jamiepinheiro.com)

      (Here’s the [raw dataset](https://top-talkers.jamiepinheiro.com/?raw=1). Each title’s underlying data-source is linked on [it’s page](https://top-talkers.jamiepinheiro.com/title/the-office) in the footer)

    2. TwoDogsInATrenchcoat on

      I find it very interesting the „other“ category in friends is so much smaller than the office.

    3. StrategyTop7612 on

      Disappointed you did Friends and the office and himym, but not seinfeld.

    Leave A Reply