AvishaiTsabari/Tokenization-Stemming-Lemmatization
0
Tokenization, Stemming & Lemmatization
Below is an interactive Gradio app demonstrating three core text‐processing steps. Use the Analyze button to see how each method transforms text.
Example Text
I ran yesterday on the beach. I love running during sunset!When to use which method
- Tokenization Splits raw text into atomic units (words, subwords, punctuation). ✅ Use when you need to feed text into a model (e.g. HF Transformers), count word frequencies, or do any indexing/lookup on the smallest building blocks.
- Stemming Heuristic chopping of word endings to reduce to a “stem” (e.g.
running→run,runs→run, but alsobetter→better). ✅ Use when you need a very fast, rule‐based normalization (e.g. search indexing), and you can tolerate over‐ or under‐stemming.
- Lemmatization Returns the dictionary (lemma) form of a word using vocabulary and POS information (e.g.
running→run,better→good). ✅ Use when you need linguistically accurate normalization for downstream tasks (POS tagging, sentiment analysis, etc.).
Conclusions
- NLTK’s default lemmatizer treats every token as a noun, so:
WordNetLemmatizer().lemmatize("running") # → "running" unless you explicitly pass pos="v" or perform POS‐aware tagging.
- spaCy’s built‐in pipeline automatically tags and lemmatizes, so:
nlp("I love running")[2].lemma_ # → "run"out of the box.
Key takeaway: Without POS‐awareness, “running” remains “running.” With a POS‐aware lemmatizer (NLTK + tagger, or spaCy), it normalizes to “run.”
