PyPDF is unmaintained and abandoned (as noted on its home page
http://pybrary.net/pyPdf/) and was never updated to Python 3. PyPDF2 is
a fork which provides a mostly compatible API and is P3-compatible.
Replace PyPDF by PyPDF2.
Commit 3ced0ff61 removed the support of Microsoft documents for
indexation. It makes sense for the old formats such as '.doc' since it
requires an external tool ('antiword'), which could lead to a security
issue. However, the new formats such '.docx' are simple xml files,
therefore they could be indexed with the usual XML parsing tools.
opw-677235