[FIX] document: PDF indexing

Backport from 12.0, commit : 1b753b0d53
note: there was also report of PDF content blocking that when indexed
blocked an instance worker indefinitely

Since PyPDF2 gives bad results for PDF indexing, stop using it since it
raises more issues than it helps users.

opw-2044679

closes odoo/odoo#35310

closes odoo/odoo#35525

Signed-off-by: Jorge Pinna Puissant (jpp) <jpp@odoo.com>
This commit is contained in:
Nicolas Martinelli
2019-08-07 11:30:51 +00:00
committed by Jorge Pinna Puissant
parent f93fcaffb5
commit dff2e242a5
+4 -1
View File
@@ -9,7 +9,7 @@ import zipfile
from odoo import api, models
_logger = logging.getLogger(__name__)
FTYPES = ['docx', 'pptx', 'xlsx', 'opendoc', 'pdf']
FTYPES = ['docx', 'pptx', 'xlsx', 'opendoc']
def textToString(element):
buff = u""
@@ -92,6 +92,9 @@ class IrAttachment(models.Model):
def _index_pdf(self, bin_data):
'''Index PDF documents'''
# extractText gives very bad results for indexing, hence we don't index PDF anymore. A
# better alternative is probably PDFMiner.six, but not for stable.
# See POC at https://github.com/odoo/odoo/pull/27568.
buf = u""
if bin_data.startswith(b'%PDF-'):
f = io.BytesIO(bin_data)