Commit Graph
13 Commits
Author SHA1 Message Date
Xavier Morel 01875541b1 [CHG] core, web: deprecate t-raw
Add a big fat warning when the qweb compiler finds a `t-raw`.

`t-esc` should now be used everywhere, the use-case for `t-raw` should
be handled by converting the corresponding values to `Markup`
objects. Even though it's convenient, this constructor *should never
be made available in the qweb rendering context* (maybe that should be
checked for explicitely?).

Replace `werkzeug.escape` by `markupsafe.escape` in
`odoo.tools.html_escape`, this means the output of `html_escape` is
markup-safe.

Updated qweb to work correctly with escaping and `Markup`, amongst
other things QWeb bodies should be markup-safe internally (so that a
`t-set` value can be fed into a `t-esc`). See at the bottom for the
attributes handling as it's a bit complicated.

`to_text` needed updating: `markupsafe.Markup` is a subclass of `str`,
but `str` is not a passthrough for strings. So `Markup` instances
going through would be converted to normal `str`, losing their safety
flag. Since qweb internally uses `to_text` on pretty much
everything (in order to handle None / False), this would then cause
almost every `Markup` to get mistakenly double-escaped.

Also mark a bunch of APIs as markup-safe by default

* html_sanitize output.
* HTML fields content, sanitization is applied on intake (so stripped
  by the trip through the database) and if the field is unsanitised
  the injection is very much intentional, probably. Note: this
  includes automatically decoding bytes as a number of default values
  & computes yield bytes, which Markup will happily accept... by
  repr-ing them which is useless. This is hard to notice without `-b`.
* Script-safe json, it's rather the point (though it uses a
  non-standard escaping scheme).
* Note that `nl2br`, kinda: it should work correctly whether or not
  the input is markup-safe, this means we should not need to escape
  values fed to `nl2br`, but it doesn't hurt either.

Update some qweb field serialisations to mark their output as
markup-safe when necessary (e.g. monetary, barcode,
contact). Otherwise either using proper escaping internally or doing
nothing should do the trick.

Also update qweb to return markup-safe bytes: we want qweb to return
markup-safe contents as a common use-case is to render something with
one template, and inject its content in an other one (with Python code
inbetween, as `t-call` works a bit differently and does not go through
the external rendering interface).

However qweb returns `bytes` while `Markup` extends `str`. After a
quick experiment with changing qweb rendering to return `str` (rather
unmitigated failure I fear), it looks like the safest tack is to add a
somewhat similar bytes-based type, which decodes to a `Markup` but
keeps to bytes semantics.

For debugging and convenience reasons, MarkupSafeBytes does *not*
stringify and raises an error instead (`__repr__` works fine). This is
to avoid implicit stringifications which do the wrong thing (namely
create a string `"b'foo'"`).

Also add some configuration around BytesWarning (which still has to be
enabled at the interpreter level via `-b`, there's no way to enable it
programmatically smh), and monkeypatch `showwarning` to show warning
tracebacks, as it's common for warnings to be triggered in the bowels
of the application, and hard to relate to business logic without the
complete traceback.

`t-out`
=======

`t-esc` is a bit confusing for the new behaviour of "maybe escape
maybe not", so add a `t-out` alias with the same behaviour.

Unlike `t-raw`, `t-esc` is only soft-deprecated for now: there are
thousands of instances, so editing all the templates is not
great. Eventually we'll add a `ci/style` to prevent addition of new
ones, and eventually we might do a bulk-replace and hard-deprecate.

Attributes handling
===================

There are a few issues with respect to attributes. The first issue is
that markup-safe content is not necessarily attributes-safe
e.g. markup-safe content can contain unescaped `<` or double-quotes
while attributes can not. So we must forcefully escape the input, even
if it's supposedly markup-safe already.

This causes a problem for script-safe JSON: it's markup-safe but
really does its own thing. So instead of escaping it up-front and
wrapping it in Markup, make script-safe JSON its own type which
applies JSON-escaping *during the `__html__` call.

This way if a script-safe JSON object goes through `markupsafe.escape`
we'll apply script-safe escaping, otherwise it'll be treated as a
regular strings and eventually escaped the normal way.

A second issue was the processing of format-valued
attributes (`t-attf`): literal segments should always be markup-safe,
while non-literal may or may not be. This turns out to be an issue if
the non-literal segment *is* markup-safe: in that case when the
literal and non-literal segments get concatenated the literal segments
will get escaped, then attributes serialization will escape
them *again* leading to doubly-escaped content in attributes.

The most visible instance of this was the `snippet_options` template,
specifically:

    <t t-set="so_content_addition_selector" t-translation="off">blockquote, ...</t>
    <div id="so_content_addition"
        t-att-data-selector="so_content_addition_selector"
        t-attf-data-drop-near="p, h1, h2, h3, .row > div > img, #{so_content_addition_selector}"
        data-drop-in=".content, nav"/>

Here `so_content_addition_selector` is a qweb body therefore
markup-safe, When concatenated with the literal part of
`t-atff-data-drop-near` it would cause the HTML-escaping of that
yielding a new Markup object. Normal attributes processing would then
strip the markup flag (using `str()`) and escape it again, leading to
doubly-escaped literals.

The original hack around was to unescape() `Markup` content before
stringifying it and escaping it again, in the attribute serialization
method (`_append_attributes`).

That's pretty disgusting, after some more consideration & testing it
looks like a much better and safer fix is to ensure the
expression (non-literal) segments of format strings always result in
`str`, never `Markup`, which is easy enough: just all `str()` on the
output of strexpr. We could also have concatenated all the bits using
`''.join` instead of repeated concatenation (`+`).

Also add a check on the type of the format string for safety, I think
it should always be a proper str and the bytes thing is only when
running in py2 (where lxml uses bytestrings as a space optimization
for ascii-only values) but it should not hurt too much to perform a
single typecheck assertion on the value... instead of performing one
per literal segment.

Note: we may need to implement unescape anyway, because it's still
possible to get double-escaping with the current scheme: given an
explicitly escape-ed `foo` and `t-att-foo="foo"`, `foo` will be
re-escaped.

fixup! [CHG] core, web: deprecate t-raw
2021-04-29 05:34:19 +00:00
Xavier Morel badb95fbce [FIX] core: further pycompat cleanup
odoo/odoo#28519 removed large parts of pycompat, but left reraise
despite that not having much value.

Remove that helper and replace it by just a `raise` in most cases:
when raising from an except block, the old exception is automatically
chained to the new one, no need to mess around.

There is one exception: in http we have to re-raise an existing
exception explicitly (aka `raise exc` rather than just `raise).

This is less than ideal as Python *concatenates* stacks: the
previously reified stack (from the except clause) is stacked on top of
the new stack (from this raises), this leads to tracebacks "jumping
around" at the break point of the handler and is somewhat confusing.

So we want to use explicit chaining (`raise a from b`) with the
"source" providing the caught exception's original traceback and the
child providing the rest.

However since callers rely on the exception making sense, we need the
re-raised exception to be the original[0]. Copying the exception
doesn't work (see [0]), chaining an exception to itself doesn't
do anything useful, and while we could probably copy exceptions using
the pickle method[1] that's still risky.

So the most reliable option seems to be to create a new "cause"
exception, move the old traceback over to it, then re-raise the
original exception having cleared its traceback, chained to new the
cause.

[0] or a copy thereof but Odoo exceptions don't all work properly with
    copy.copy and we don't want that to fail so not really an option,
    we can't rely / bet on every new exception being cleanly copy-able
[1] create an "empty" instance using __new__ (or an instance of
    something else onto which we re-set the __class__ in case the exctype
    actually overrides __new__) then copy the __dict__

closes odoo/odoo#39709

Signed-off-by: Xavier Morel (xmo) <xmo@odoo.com>
2019-11-05 08:14:10 +00:00
Adrian Torres 758382b3a7 [REM] pycompat: remove python 2 shims and helpers
Odoo no longer supports python 2, thus some of these helpers can and
have been replaced by python 3 built-ins, therefore there is no need for
them to stay defined.

The removed helpers are:
    * izip, imap and ifilter
    * unichr, text_type
    * implements_to_string, implements_iterator
    * string_types, integer_types
    * to_native

The python 2 shims have also been removed, and only the python 3 helpers
have been kept, because they can still be usable (i.e. accepting
both bytes and str for functions that can only accept one of the two)

[REM] pyjsparser: remove PY3 shims

They're no longer necessary as Odoo doesn't officially support python 2
anymore.

closes odoo/odoo#28519
2018-11-29 09:28:17 +00:00
Raphael Collet 6a601fe6dc [IMP] test-pylint: check for undefined variables 2018-05-16 13:58:39 +02:00
Xavier Morel 4c3846e7e0 [FIX] tools: use codecs for csv reader/writer
The class PoFile was already using it.
TextIOWrapper closes its underlying buffer causing the "Synchronize Terms" action
to fail (trying to read on a closed buffer).

Fixes #19911
2017-10-11 11:42:31 +02:00
Olivier Dony 695716efb0 [FIX] P3: remove pycompat.{keys,items,values} helpers
Now that we're closer to switching to P3 for good, these helpers have
outlived their usefulness, and mostly add noise.

All remaining dict.iter*() or dict.view*() must be converted to the
normal keys(), values() or items() calls.

Whenever the result is likely to be used for more than the scope of a
loop, or when the dict needs to be modified during iteration, the calls
must be wrapped in a ``list()``, to protect the new P3 semantics.
Those cases are very exceptional.

Also removed some dead code or improved the API to remove unnecessary
conversions.
2017-08-20 23:25:54 +02:00
Xavier Morel be7c5aefdf [FIX] P3: CSV reading & writing 2017-08-20 23:25:54 +02:00
Xavier Morel 7dd062f835 [FIX] P3: text model types
* remove references to basestring & unicode (use relevant pycompat
  helpers)
* remove some str calls (either entirely or replaced by relevant
  helper, either text or native)
* use better API to avoid unnecessary conversions
* remove some XML declarations in views
2017-08-20 23:25:54 +02:00
Xavier Morel 3824b5dcc1 [FIX] P3: fix base64 and StringIO uses
* StringIO removed from stdlib, replace with io
* try to correctly handle BytesIO/StringIO (one is for bytes the other
  is for text)
* fix base64: Python 3 removed bytes-encoding and bytes-bytes
  codecs (via #encode) so replace all calls to str.encode('base64'),
  also b64encode is a bytes->bytes conversion so attempt to properly
  handle that

issue #8530
2017-08-20 23:25:54 +02:00
Xavier Morel 3dd3790597 [FIX] P3: raise exception with existing traceback
In Python 2, to raise a new exception but reuse a traceback requires a
special form of ``raise`` (``raise etype, evalue, tb``).

In Python 3, this is now done via a ``with_traceback`` method on
exception objects.

However this requires a bit of trickery as the former is invalid
syntax in Python 3, hence pycompat bridge created via an exec for
Python 2.
2017-05-12 16:15:39 +02:00
xmo-odoo fffaf735f5 [FIX] P3: list -> iterable builtins (#16811)
In Python 3:

* various builtins and dict methods were changed to return
  view/iterable objects rather than lists
* and the separate Python 2 view/iterable builtins and methods were
  removed altogether

This is problematic when using these items as list (which the happens
repeatedly in Odoo), but more viciously when iterating *multiple times*
over them (which also happens, which I've messed up multiple times while
writing this, and which is a pain to debug even when you've just created
the issue).

Convert all code using these to semantics-matching cross-version
helper functions to get the LCD behaviour between P2 and P3, and
forbid the builtins via lint.

issue #8530
2017-05-10 09:39:55 +02:00
xmo-odoo fdaf967bb0 [FIX] forgot a bit of version checking 2017-04-27 15:03:28 +02:00
xmo-odoo 2e6a589f41 [FIX] builtins removed from Python 3
* Reverse wrapper courtesy of @rco-odoo's original P3 branch
* thin compat module stripped down from werkzeug (to augment as needed)

issue 8530
2017-04-27 13:59:33 +02:00