lxml.HTMLParser detect the content encoding by itself if not specified.
in the issue leading to this fix, content in a mail with u'\xa4'
( ) was transformed intto u'\xc2\xa4' (Â ) since HTMLParser
selected the wrong encoding (latin-1 instead of utf-8 for example).
this commit specify the known encoding and revert 3d00c210f which was
another attempt at this issue.