Welcome to OGeek Q&A Community for programmer and developer-Open, Learning and Share
Welcome To Ask or Share your Answers For Others

Categories

0 votes
311 views
in Technique[技术] by (71.8m points)

python - Reading nothing from a pdf file using PyPDF2

I want to convert a pdf file to json format. So I was using PyPDF2 module to read the pdf. But I am unable to read it. I gives me some " " characters but no text. The pdf I am using can be retrieve from here: pdf_to_json.pdf

The code I am using is:

import PyPDF2

file = open("pdf_to_json.pdf", "rb")

pdf = PyPDF2.PdfFileReader(file)

page_one = pdf.getPage(0)

page_one.extractText()

It's returning something like this:

'














































































'

DISCLAIMER: The pdf is in spanish

question from:https://stackoverflow.com/questions/65649347/reading-nothing-from-a-pdf-file-using-pypdf2

与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…
Welcome To Ask or Share your Answers For Others

1 Reply

0 votes
by (71.8m points)
Waitting for answers

与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…
OGeek|极客中国-欢迎来到极客的世界,一个免费开放的程序员编程交流平台!开放,进步,分享!让技术改变生活,让极客改变未来! Welcome to OGeek Q&A Community for programmer and developer-Open, Learning and Share
Click Here to Ask a Question

...