budou icon indicating copy to clipboard operation
budou copied to clipboard

Non-breaking space character (/u00A0) causes AssertionError

Open lacymorrow opened this issue 3 years ago • 0 comments

Here is the problem string: Chatbot\u00a0\u2013

Traceback (most recent call last):
  File "<console>", line 5, in <module>
  File "/usr/local/lib/python3.6/site-packages/budou/parser.py", line 78, in parse
    chunks = self.segmenter.segment(source, language)
  File "/usr/local/lib/python3.6/site-packages/budou/tinysegmentersegmenter.py", line 94, in segment
    assert source[seek] == ' '
AssertionError

https://github.com/google/budou/blob/87d9b81bdd21d1a41436df140e1bc08d817119a3/budou/tinysegmentersegmenter.py#L94

lacymorrow avatar Mar 04 '21 22:03 lacymorrow