budou
budou copied to clipboard
Non-breaking space character (/u00A0) causes AssertionError
Here is the problem string: Chatbot\u00a0\u2013
Traceback (most recent call last):
File "<console>", line 5, in <module>
File "/usr/local/lib/python3.6/site-packages/budou/parser.py", line 78, in parse
chunks = self.segmenter.segment(source, language)
File "/usr/local/lib/python3.6/site-packages/budou/tinysegmentersegmenter.py", line 94, in segment
assert source[seek] == ' '
AssertionError
https://github.com/google/budou/blob/87d9b81bdd21d1a41436df140e1bc08d817119a3/budou/tinysegmentersegmenter.py#L94