python - unicode().decode('utf-8', 'ignore') raising UnicodeEncodeError

Question

Welcome To Ask or Share your Answers For Others

python - unicode().decode('utf-8', 'ignore') raising UnicodeEncodeError

posted Oct 17, 2021 in Technique[技术] by 深蓝 (71.8m points)

python - unicode().decode('utf-8', 'ignore') raising UnicodeEncodeError

Here is the code:

>>> z = u'u2022'.decode('utf-8', 'ignore')
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/usr/lib/python2.6/encodings/utf_8.py", line 16, in decode
    return codecs.utf_8_decode(input, errors, True)
UnicodeEncodeError: 'latin-1' codec can't encode character u'u2022' in position 0: ordinal not in range(256)

Why is UnicodeEncodeError raised when I am using .decode?

Why is any error raised when I am using 'ignore'?

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Reply

深蓝 · Answer 1 · 2021-10-17T00:13:22+0000

When I first started messing around with python strings and unicode, It took me awhile to understand the jargon of decode and encode too, so here's my post from here that may help:

Think of decoding as what you do to go from a regular bytestring to unicode and encoding as what you do to get back from unicode. In other words:

You de - code a str to produce a unicode string

and en - code a unicode string to produce an str.

So:

unicode_char = u'xb0'

encodedchar = unicode_char.encode('utf-8')

encodedchar will contain your unicode character, displayed in the selected encoding (in this case, utf-8).

Categories

python - unicode().decode('utf-8', 'ignore') raising UnicodeEncodeError

python - unicode().decode('utf-8', 'ignore') raising UnicodeEncodeError

Please log in or register to add a comment.

Please log in or register to reply this article.

1 Reply

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags