`markupsafe.escape` always escapes single and double quotes, and
escapes them to their numeric values rather than symbolic
According to pallets/jinja@f35e28154f,
this is for compatibility with HTML 3.2: the only named entities in
the HTML 3.2 DTD are `amp`, `gt`, and `lt`.
Update tests to match.
Add a big fat warning when the qweb compiler finds a `t-raw`.
`t-esc` should now be used everywhere, the use-case for `t-raw` should
be handled by converting the corresponding values to `Markup`
objects. Even though it's convenient, this constructor *should never
be made available in the qweb rendering context* (maybe that should be
checked for explicitely?).
Replace `werkzeug.escape` by `markupsafe.escape` in
`odoo.tools.html_escape`, this means the output of `html_escape` is
markup-safe.
Updated qweb to work correctly with escaping and `Markup`, amongst
other things QWeb bodies should be markup-safe internally (so that a
`t-set` value can be fed into a `t-esc`). See at the bottom for the
attributes handling as it's a bit complicated.
`to_text` needed updating: `markupsafe.Markup` is a subclass of `str`,
but `str` is not a passthrough for strings. So `Markup` instances
going through would be converted to normal `str`, losing their safety
flag. Since qweb internally uses `to_text` on pretty much
everything (in order to handle None / False), this would then cause
almost every `Markup` to get mistakenly double-escaped.
Also mark a bunch of APIs as markup-safe by default
* html_sanitize output.
* HTML fields content, sanitization is applied on intake (so stripped
by the trip through the database) and if the field is unsanitised
the injection is very much intentional, probably. Note: this
includes automatically decoding bytes as a number of default values
& computes yield bytes, which Markup will happily accept... by
repr-ing them which is useless. This is hard to notice without `-b`.
* Script-safe json, it's rather the point (though it uses a
non-standard escaping scheme).
* Note that `nl2br`, kinda: it should work correctly whether or not
the input is markup-safe, this means we should not need to escape
values fed to `nl2br`, but it doesn't hurt either.
Update some qweb field serialisations to mark their output as
markup-safe when necessary (e.g. monetary, barcode,
contact). Otherwise either using proper escaping internally or doing
nothing should do the trick.
Also update qweb to return markup-safe bytes: we want qweb to return
markup-safe contents as a common use-case is to render something with
one template, and inject its content in an other one (with Python code
inbetween, as `t-call` works a bit differently and does not go through
the external rendering interface).
However qweb returns `bytes` while `Markup` extends `str`. After a
quick experiment with changing qweb rendering to return `str` (rather
unmitigated failure I fear), it looks like the safest tack is to add a
somewhat similar bytes-based type, which decodes to a `Markup` but
keeps to bytes semantics.
For debugging and convenience reasons, MarkupSafeBytes does *not*
stringify and raises an error instead (`__repr__` works fine). This is
to avoid implicit stringifications which do the wrong thing (namely
create a string `"b'foo'"`).
Also add some configuration around BytesWarning (which still has to be
enabled at the interpreter level via `-b`, there's no way to enable it
programmatically smh), and monkeypatch `showwarning` to show warning
tracebacks, as it's common for warnings to be triggered in the bowels
of the application, and hard to relate to business logic without the
complete traceback.
`t-out`
=======
`t-esc` is a bit confusing for the new behaviour of "maybe escape
maybe not", so add a `t-out` alias with the same behaviour.
Unlike `t-raw`, `t-esc` is only soft-deprecated for now: there are
thousands of instances, so editing all the templates is not
great. Eventually we'll add a `ci/style` to prevent addition of new
ones, and eventually we might do a bulk-replace and hard-deprecate.
Attributes handling
===================
There are a few issues with respect to attributes. The first issue is
that markup-safe content is not necessarily attributes-safe
e.g. markup-safe content can contain unescaped `<` or double-quotes
while attributes can not. So we must forcefully escape the input, even
if it's supposedly markup-safe already.
This causes a problem for script-safe JSON: it's markup-safe but
really does its own thing. So instead of escaping it up-front and
wrapping it in Markup, make script-safe JSON its own type which
applies JSON-escaping *during the `__html__` call.
This way if a script-safe JSON object goes through `markupsafe.escape`
we'll apply script-safe escaping, otherwise it'll be treated as a
regular strings and eventually escaped the normal way.
A second issue was the processing of format-valued
attributes (`t-attf`): literal segments should always be markup-safe,
while non-literal may or may not be. This turns out to be an issue if
the non-literal segment *is* markup-safe: in that case when the
literal and non-literal segments get concatenated the literal segments
will get escaped, then attributes serialization will escape
them *again* leading to doubly-escaped content in attributes.
The most visible instance of this was the `snippet_options` template,
specifically:
<t t-set="so_content_addition_selector" t-translation="off">blockquote, ...</t>
<div id="so_content_addition"
t-att-data-selector="so_content_addition_selector"
t-attf-data-drop-near="p, h1, h2, h3, .row > div > img, #{so_content_addition_selector}"
data-drop-in=".content, nav"/>
Here `so_content_addition_selector` is a qweb body therefore
markup-safe, When concatenated with the literal part of
`t-atff-data-drop-near` it would cause the HTML-escaping of that
yielding a new Markup object. Normal attributes processing would then
strip the markup flag (using `str()`) and escape it again, leading to
doubly-escaped literals.
The original hack around was to unescape() `Markup` content before
stringifying it and escaping it again, in the attribute serialization
method (`_append_attributes`).
That's pretty disgusting, after some more consideration & testing it
looks like a much better and safer fix is to ensure the
expression (non-literal) segments of format strings always result in
`str`, never `Markup`, which is easy enough: just all `str()` on the
output of strexpr. We could also have concatenated all the bits using
`''.join` instead of repeated concatenation (`+`).
Also add a check on the type of the format string for safety, I think
it should always be a proper str and the bytes thing is only when
running in py2 (where lxml uses bytestrings as a space optimization
for ascii-only values) but it should not hurt too much to perform a
single typecheck assertion on the value... instead of performing one
per literal segment.
Note: we may need to implement unescape anyway, because it's still
possible to get double-escaping with the current scheme: given an
explicitly escape-ed `foo` and `t-att-foo="foo"`, `foo` will be
re-escaped.
fixup! [CHG] core, web: deprecate t-raw
* from collections import <ABC> is deprecated, unclear why the
deprecation warning didn't appear before (possibly only appears in
3.7/3.8?) either way `collections.abc` should be 3.3+ so switch
everything to it.
* add some more ignores on third-party packages deprecation
warnings (meh)
* while at it, mitigate generation of non-breaking space on some
versions of Babel (in the french locale used by our tests anyway)
closesodoo/odoo#47581
Related: odoo/enterprise#9214
Signed-off-by: Xavier Morel (xmo) <xmo@odoo.com>
load_lang was a kind of hybrid method trying to active or creating a
language if not found. This was error prone.
Instead rely on two methods with clear purpose:
ResLang._create_lang(lang, lang_name=None)
- create a new res.lang entry using the locale of the server
return the res.lang record to match the API of _activate_lang
ResLang._active_lang(code)
- activate the given code lang
Most of the time, _active_lang is what is expected
tools.trans_load_data and IrTranslation._load_module_terms no longer
activate the language if not active.
Loading the translations should be explicit on an activated language,
it is too error prone to silently activate/create a language if not
found.
Remove lang_name from trans_load_data as no longer needed.
The option `digital` has been added. It allows to display "01:00" instead of "1
hour". This behaviour was implemented in a report (timesheet) but the logic
was done in the XML, which is not adequate.
This commit also allows to have a negative value for this widget.
* StringIO removed from stdlib, replace with io
* try to correctly handle BytesIO/StringIO (one is for bytes the other
is for text)
* fix base64: Python 3 removed bytes-encoding and bytes-bytes
codecs (via #encode) so replace all calls to str.encode('base64'),
also b64encode is a bytes->bytes conversion so attempt to properly
handle that
issue #8530