[FIX] tools: nbsp+sanitize = write

lxml.html.clean module we are using to sanitize html takes either a
tree/element or string, and return us an element of the same type.

We are passing it an unicode string, and it gives us back an unicode
string and if it is possible convert entities to utf-8.

This lead to an issue in html inline fields containing an ` `
entity, the entity was converted by the sanitizer to its unicode
character (U+00A0).

Then, when the field value is loaded to be displayed, the internal
value contains U+00A0 characters, but when it is inserted in the page
the browser converts them to ` `. Hence in this instance, the
internal value is never equal to the displayed value and a write is
always triggered even without editing.

This commit simple modify the sanitizer so value saved in database
contain ` `.

closes #14569
opw-689274
This commit is contained in:
Nicolas Lempereur
2016-12-06 18:54:46 +01:00
parent aec8e2582f
commit b51d21c5b8
+2
View File
@@ -223,6 +223,8 @@ def html_sanitize(src, silent=True, sanitize_tags=True, sanitize_attributes=Fals
cleaned = cleaned.replace('%7C', '|')
cleaned = cleaned.replace('&lt;%', '<%')
cleaned = cleaned.replace('%&gt;', '%>')
# html considerations so real html content match database value
cleaned.replace(u'\xa0', '&nbsp;')
except etree.ParserError, e:
if 'empty' in str(e):
return ""