Team Ai
Apppublic

AvishaiTsabari/Tokenization-Stemming-Lemmatization

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes
App README

Tokenization, Stemming & Lemmatization

Below is an interactive Gradio app demonstrating three core text‐processing steps. Use the Analyze button to see how each method transforms text.


Example Text

text
I ran yesterday on the beach. I love running during sunset!

When to use which method

  • —Tokenization Splits raw text into atomic units (words, subwords, punctuation). ✅ Use when you need to feed text into a model (e.g. HF Transformers), count word frequencies, or do any indexing/lookup on the smallest building blocks.
  • —Stemming Heuristic chopping of word endings to reduce to a “stem” (e.g. running → run, runs → run, but also better → better). ✅ Use when you need a very fast, rule‐based normalization (e.g. search indexing), and you can tolerate over‐ or under‐stemming.
  • —Lemmatization Returns the dictionary (lemma) form of a word using vocabulary and POS information (e.g. running → run, better → good). ✅ Use when you need linguistically accurate normalization for downstream tasks (POS tagging, sentiment analysis, etc.).

Conclusions

  • —NLTK’s default lemmatizer treats every token as a noun, so:
python
  WordNetLemmatizer().lemmatize("running")  # → "running"

unless you explicitly pass pos="v" or perform POS‐aware tagging.

  • —spaCy’s built‐in pipeline automatically tags and lemmatizes, so:
python
  nlp("I love running")[2].lemma_  # → "run"

out of the box.

Key takeaway: Without POS‐awareness, “running” remains “running.” With a POS‐aware lemmatizer (NLTK + tagger, or spaCy), it normalizes to “run.”