Some distributions still bundle chardet 2.3, which have some guessing
divergences / incompatibilities with python (resolved in 3.x):
1. UTF-{16,32} with BOM is guessed as LE/BE, which when used to decode
the string doesn't strip out the BOM. Handle this by checking if
the BOM is present and converting the encoding name to the
non-marked version in that case.
2. The ISO-8859-1 test string is guessed as ISO-8859-2 (TBF the
decoding does make some sense). Allow multiple targets/guesses to
"fix" that.
closesodoo/odoo#33179
Signed-off-by: Xavier Morel (xmo) <xmo@odoo.com>
* expand auto-detected date and time patterns (e.g. %b, %I, ...)
* try to make date-pattern-detection clearer
* add a select2 dropdown for date patterns (with a bunch of
preselected patterns) rather than just an input
* also try to improve other column-matching bits (e.g. less reliance
on exceptions, attempts to avoid redundant work)
It would probably be even better to iterate the file content and get
the non-quoted non-alphanumeric characters as separator
candidates (instead of a hard-coded list) however Python does not seem
to have a decoding iterator (taking bytes and yielding an iterator of
codepoints or even grapheme clusters) — incidentally uniseg seems to
require up-front decoding as well — so that's not really convenient as
we may be dealing with large-ish files and not want to load it
entirely in memory.
An alternative would be to use TextIOWrapper and iterate the file by
buffers of a few ks, and classify that based on either codepoints or
grapheme clusters.
* if an encoding is explicitly specified, use it and don't guess
* otherwise guess and return the guessed encoding so it can be
displayed in the configuration UI
* fix less-than-stellar configuration & behaviour of select2 inputs
to properly reflect underlying values as they get modified, to
correctly handle future configuration guesses