`get_link_preview_from_html` was coded to avoid fetching the entire page content when it only needs page data, however it tries to decode the entire input buffer even though the fetch window might cut off a codepoint in two.
Fix by stripping out anything which follows the `</head>` tag, which might contain partial codepoints.
Also add a few more improvements:
- avoid visiting the entire buffer if we parsed more than 16k (as unlikely as that is), even when offsetted `find` returns an offset from the start of the buffer
- don't decode upfront, as `html.fromstring` will decode just fine internally (better really as it has decoding fallbacks)
closesodoo/odoo#128240
Signed-off-by: Xavier Morel (xmo) <xmo@odoo.com>
Improve the _get_link_preview_from_url method to:
- Only load necessary data (only the <head> instead of
the whole page).
- Handle optional request session.
- Handle all image mimetypes.
- Add a fallback on the <title> tag when no og:title have
been found.
Moving the method out of the model to facilitate
its use as a tool.
Task-3234864
Part-of: odoo/odoo#122087