Discussion about this post

User's avatar
Natt S.'s avatar

Thank you. Identity resolution and consolidation is super important and not talked about enough considering its wide surface area for potentially occurring in any organization's records.

Deniz's avatar

Sonal, great write up. I'm currently going through the process of data cleansing and NER then clustering. Seeing similar issues and pain-points. Aside from clustering and de-duplication, I found that extraction from raw formats, especially PDF, adds to this pain. Any opinino or techniques you'd suggest from accurate PDF extractions? Curious, I have a whole pipeline to overcome the challenge through a layered approach but wanted to hear your opinion too.

5 more comments...

No posts

Ready for more?