Welcome to OGeek Q&A Community for programmer and developer-Open, Learning and Share
Welcome To Ask or Share your Answers For Others

Categories

0 votes
140 views
in Technique[技术] by (71.8m points)

html - How can I copy the content of a website to a Markdown file with Python?

I want to write a Python function that takes an URL as an argument and outputs a Markdown file with the content of the webpage. The images embedded in the website should be download and appropriately referenced in the Markdown file.

I wrote this code

import requests
import html2text

# The URL
link = "https://www.some.website"
f = requests.get(link)

# URL content to plain text (HTML)
textHtml = f.text

# HTML text to MD text
h = html2text.HTML2Text()
textMd = h.handle(textHtml)

# MD text is written to file
text_file = open("output.md", "w")
text_file.write(textMd)
text_file.close()

I think it does the job in downloading the text and formatting it into a Markdown file but I don't know how to download images and add references in the Markdown file to the local image files.

How can I do this?

Thanks in advance!

question from:https://stackoverflow.com/questions/65901004/how-can-i-copy-the-content-of-a-website-to-a-markdown-file-with-python

与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…
Welcome To Ask or Share your Answers For Others

1 Reply

0 votes
by (71.8m points)
Waitting for answers

与恶龙缠斗过久,自身亦成为恶龙;凝视深渊过久,深渊将回以凝视…
OGeek|极客中国-欢迎来到极客的世界,一个免费开放的程序员编程交流平台!开放,进步,分享!让技术改变生活,让极客改变未来! Welcome to OGeek Q&A Community for programmer and developer-Open, Learning and Share
Click Here to Ask a Question

...