Welcome To Ask or Share your Answers For Others

regex - How to break up a paragraph by sentences in Python

Welcome To Ask or Share your Answers For Others

1 Reply

replyed Oct 24, 2021 by 深蓝 (71.8m points)

The nltk.tokenize module is designed for this and handles edge cases. For example:

>>> from nltk import tokenize
>>> p = "Good morning Dr. Adams. The patient is waiting for you in room number 3."
>>> tokenize.sent_tokenize(p)
['Good morning Dr. Adams.', 'The patient is waiting for you in room number 3.']

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

...