String similarity metrics in Python

Question 1

String similarity metrics in Python

python algorithm string levenshtein-distance

agiliq · Sep 24, 2009 · Viewed 46.2k times · Source

Answer

Answer

I realize it's not the same thing, but this is close enough:

>>> import difflib
>>> a = 'Hello, All you people'
>>> b = 'hello, all You peopl'
>>> seq=difflib.SequenceMatcher(a=a.lower(), b=b.lower())
>>> seq.ratio()
0.97560975609756095

You can make this as a function

def similar(seq1, seq2):
    return difflib.SequenceMatcher(a=seq1.lower(), b=seq2.lower()).ratio() > 0.9

>>> similar(a, b)
True
>>> similar('Hello, world', 'Hi, world')
False

Question 2

I want to find string similarity between two strings. This page has examples of some of them. Python has an implemnetation of Levenshtein algorithm. Is there a better algorithm, (and hopefully a python library), under these contraints.

I want to do fuzzy matches between strings. eg matches('Hello, All you people', 'hello, all You peopl') should return True
False negatives are acceptable, False positives, except in extremely rare cases are not.
This is done in a non realtime setting, so speed is not (much) of concern.
[Edit] I am comparing multi word strings.

Would something other than Levenshtein distance(or Levenshtein ratio) be a better algorithm for my case?

String similarity metrics in Python

Answer

Related questions