[FIX] mail: well sanitize the mails coming from outlook desktop
Steps: - Create a ticket with a customer ( it will the email) - Answer to the mail with Outlook Desktop( a multiple lines break) - Look at the response in the discuss frame. Issue: In the frame there is much more lines break than in the original mail Cause: The format of outlook desktop mail before sanitizing looks like this ``` <div class="WordSection1"> <p class="MsoNormal">Test<o:p></o:p></p> <p class="MsoNormal"><o:p> </o:p></p> <p class="MsoNormal"><o:p> </o:p></p> <p class="MsoNormal">Two break lines<o:p></o:p></p> <p class="MsoNormal"><o:p> </o:p></p> <div> ``` So when parsing it transforms the ``<o:p></o:p>`` in ``<p></p>`` which adds more lines break when displayed. Solution: Remove empty ``<o:...>`` and ``</o:...>`` tags which are specific to outlook desktop before. opw-3089550 closes odoo/odoo#144931 X-original-commit: a3e9db87277771c9801f24a5f52aa73b98a0d879 Signed-off-by: Thibault Delavallee (tde) <tde@openerp.com>
This commit is contained in:
parent
1f42c0f468
commit
870acffba6
@@ -119,6 +119,46 @@ class TestSanitizer(BaseCase):
|
||||
for attr in ['javascript']:
|
||||
self.assertNotIn(attr, sanitized_html, 'html_sanitize did not remove enough unwanted attributes')
|
||||
|
||||
def test_outlook_mail_sanitize(self):
|
||||
case = """<div class="WordSection1">
|
||||
<p class="MsoNormal">Here is a test mail<o:p></o:p></p>
|
||||
<p class="MsoNormal"><o:p> </o:p></p>
|
||||
<p class="MsoNormal">With a break line<o:p></o:p></p>
|
||||
<p class="MsoNormal"><o:p> </o:p></p>
|
||||
<p class="MsoNormal"><o:p> </o:p></p>
|
||||
<p class="MsoNormal">Then two<o:p></o:p></p>
|
||||
<p class="MsoNormal"><o:p> </o:p></p>
|
||||
<div>
|
||||
<div style="border:none;border-top:solid #E1E1E1 1.0pt;padding:3.0pt 0in 0in 0in">
|
||||
<p class="MsoNormal"><b>From:</b> Mitchell Admin <dummy@example.com>
|
||||
<br>
|
||||
<b>Sent:</b> Monday, November 20, 2023 8:34 AM<br>
|
||||
<b>To:</b> test user <dummy@example.com><br>
|
||||
<b>Subject:</b> test (#23)<o:p></o:p></p>
|
||||
</div>
|
||||
</div>"""
|
||||
|
||||
expected = """<div class="WordSection1">
|
||||
<p class="MsoNormal">Here is a test mail</p>
|
||||
<p class="MsoNormal"> </p>
|
||||
<p class="MsoNormal">With a break line</p>
|
||||
<p class="MsoNormal"> </p>
|
||||
<p class="MsoNormal"> </p>
|
||||
<p class="MsoNormal">Then two</p>
|
||||
<p class="MsoNormal"> </p>
|
||||
<div>
|
||||
<div style="border:none;border-top:solid #E1E1E1 1.0pt;padding:3.0pt 0in 0in 0in">
|
||||
<p class="MsoNormal"><b>From:</b> Mitchell Admin <dummy@example.com>
|
||||
<br>
|
||||
<b>Sent:</b> Monday, November 20, 2023 8:34 AM<br>
|
||||
<b>To:</b> test user <dummy@example.com><br>
|
||||
<b>Subject:</b> test (#23)</p>
|
||||
</div>
|
||||
</div></div>"""
|
||||
|
||||
result = html_sanitize(case)
|
||||
self.assertEqual(result, expected)
|
||||
|
||||
def test_sanitize_unescape_emails(self):
|
||||
not_emails = [
|
||||
'<blockquote cite="mid:CAEJSRZvWvud8c6Qp=wfNG6O1+wK3i_jb33qVrF7XyrgPNjnyUA@mail.gmail.com" type="cite">cat</blockquote>',
|
||||
|
||||
@@ -202,6 +202,9 @@ def html_normalize(src, filter_callback=None):
|
||||
try:
|
||||
src = src.replace('--!>', '-->')
|
||||
src = re.sub(r'(<!-->|<!--->)', '<!-- -->', src)
|
||||
# On the specific case of Outlook desktop it adds unnecessary '<o:.*></o:.*>' tags which are parsed
|
||||
# in '<p></p>' which may alter the appearance (eg. spacing) of the mail body
|
||||
src = re.sub(r'</?o:.*?>', '', src)
|
||||
doc = html.fromstring(src)
|
||||
except etree.ParserError as e:
|
||||
# HTML comment only string, whitespace only..
|
||||
|
||||
Reference in New Issue
Block a user