It would probably be even better to iterate the file content and get
the non-quoted non-alphanumeric characters as separator
candidates (instead of a hard-coded list) however Python does not seem
to have a decoding iterator (taking bytes and yielding an iterator of
codepoints or even grapheme clusters) — incidentally uniseg seems to
require up-front decoding as well — so that's not really convenient as
we may be dealing with large-ish files and not want to load it
entirely in memory.
An alternative would be to use TextIOWrapper and iterate the file by
buffers of a few ks, and classify that based on either codepoints or
grapheme clusters.