Picture shows the iPhone Duo, a new iPhone with a foldable screen, enabling both the person holding the phone and an implied observer just barely visible on the right of the frame to view the screen, thereby making the depicted scene a 3 person interaction.
Now, there are other interactions involving three persons, some of them living under the umbrella term cuckold, hence making the device’s foldable screen the cuck screen.

If the subtitles are actual subtitles and not dubs, you can extract timestamps from the audio with a model such as whisper + timestamp support. You could then optimize the overlap of timestamps in the subtitle format (e. g. SRT) with the timestamps provided by the model. That would give you a best fit delay, potentially even shifting over time.
Given that you’re in a browser and don’t even have access to the subtitle data, that’s a purely theoretical exercise. While it might be possible to retrieve the subtitles from the site / API the browser is using and then set the delay in the browser component as well, that’d be far too much effort vs. switching to a better playback method.
You could, however, compute the best match and set that manually, if that’s all you’re after, i. e. skip the automation part. You’d still need to access the source audio and subtitles.
If the audio is in a different language, all bets are off, there is nothing to sync against. Translations also won’t help, since there are multiple ways to translate any given source material, and the origin and length don’t match up between languages, naturally.