Files
odoo_source/odoo/tools/json/scriptsafe.py
T
Xavier Morel 05746b7068 [IMP] *: script-safe JSON serialisation
* add a json_scriptsafe module, which re-exports the stdlib's json but
provides a script-safe `dumps` function, because of the way it works
the script-safe version is fine / safe to use outside script
contexts, so there's no need to provide different functions to qweb
contexts
* make the json.dumps in qweb contexts <script>-safe (should be
rendered using t-raw)
* update callsites to properly dump json:
- dump in template, not in the python code
- dump the entire structure at once, not bits and bobs inside a JSON
literal

Task 1999158

closes odoo/odoo#33682

Signed-off-by: Xavier Morel (xmo) <xmo@odoo.com>
2019-06-19 13:47:35 +00:00

51 lines
1.7 KiB
Python

# -*- coding: utf-8 -*-
import json as json_
import re
from json import *
__all__ = ['dumps', 'loads']
JSON_SCRIPTSAFE_MAPPER = {
'&': r'\u0026',
'<': r'\u003c',
'>': r'\u003e',
'\u2028': r'\u2028',
'\u2029': r'\u2029'
}
def dumps(*args, **kwargs):
""" JSON used as JS in HTML (script tags) is problematic: <script>
tags are a special context which only waits for </script> but doesn't
interpret anything else, this means standard htmlescaping does not
work (it breaks double quotes, and e.g. `<` will become `&lt;` *in
the resulting JSON/JS* not just inside the page).
However, failing to escape embedded json means the json strings could
contains `</script>` and thus become XSS vector.
The solution turns out to be very simple: use JSON-level unicode
escapes for HTML-unsafe characters (e.g. "<" -> "\u003C". This removes
the XSS issue without breaking the json, and there is no difference to
the end result once it's been parsed back from JSON. So it will work
properly even for HTML attributes or raw text.
Also handle U+2028 and U+2029 the same way just in case as these are
interpreted as newlines in javascript but not in JSON, which could
lead to oddities and issues.
.. warning::
except inside <script> elements, this should be escaped following
the normal rules of the containing format
Cf https://code.djangoproject.com/ticket/17419#comment:27
"""
# replacement can be done straight in the serialised JSON as the
# problematic characters are not JSON metacharacters (and can thus
# only occur in strings)
return re.sub(
r'[<>&\u2028\u2029]',
lambda m: JSON_SCRIPTSAFE_MAPPER[m[0]],
json_.dumps(*args, **kwargs),
)