* remove references to basestring & unicode (use relevant pycompat
helpers)
* remove some str calls (either entirely or replaced by relevant
helper, either text or native)
* use better API to avoid unnecessary conversions
* remove some XML declarations in views
* StringIO removed from stdlib, replace with io
* try to correctly handle BytesIO/StringIO (one is for bytes the other
is for text)
* fix base64: Python 3 removed bytes-encoding and bytes-bytes
codecs (via #encode) so replace all calls to str.encode('base64'),
also b64encode is a bytes->bytes conversion so attempt to properly
handle that
issue #8530
PyPDF is unmaintained and abandoned (as noted on its home page
http://pybrary.net/pyPdf/) and was never updated to Python 3. PyPDF2 is
a fork which provides a mostly compatible API and is P3-compatible.
Replace PyPDF by PyPDF2.
Commit 3ced0ff61 removed the support of Microsoft documents for
indexation. It makes sense for the old formats such as '.doc' since it
requires an external tool ('antiword'), which could lead to a security
issue. However, the new formats such '.docx' are simple xml files,
therefore they could be indexed with the usual XML parsing tools.
opw-677235