Get in Touch

Course Outline

Comprehensive training outline

  1. Introduction to NLP
    • Concepts of NLP
    • Popular NLP frameworks
    • Commercial uses of NLP
    • Web data scraping techniques
    • Utilising APIs to fetch text data
    • Managing and storing text corpora with relevant metadata
    • Benefits of Python and an NLTK quick start
  2. Practical Understanding of a Corpus and Dataset
    • The necessity of a corpus
    • Corpus analysis methods
    • Varieties of data attributes
    • File formats suitable for corpora
    • Dataset preparation for NLP tasks
  3. Understanding the Structure of Sentences
    • Core NLP components
    • Natural language understanding
    • Morphological analysis: stemming, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Managing ambiguity
  4. Text data preprocessing
    • Raw text corpus
      • Sentence tokenization
      • Stemming raw text
      • Lemmatizing raw text
      • Removal of stop words
    • Raw sentence corpus
      • Word tokenization
      • Word lemmatization
    • Handling Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Customized and practical preprocessing strategies
  5. Analyzing Text data
    • Fundamental NLP features
      • Parsers and parsing techniques
      • POS tagging and taggers
      • Name entity recognition
      • N-grams
      • Bag of words approach
    • Statistical aspects of NLP
      • Linear algebra concepts for NLP
      • Probabilistic theory in NLP
      • TF-IDF
      • Vectorization techniques
      • Encoders and Decoders
      • Data normalization
      • Probabilistic Models
    • Advanced feature engineering in NLP
      • Foundations of word2vec
      • Structural components of the word2vec model
      • Operational logic of word2vec
      • Extensions of the word2vec concept
      • Practical applications of the word2vec model
    • Case study: Bag of words application in automatic text summarization using simplified and true Luhn's algorithms
  6. Document Clustering, Classification and Topic Modeling
    • Document clustering and pattern discovery (including hierarchical and k-means clustering)
    • Document comparison and classification using TFIDF, Jaccard, and cosine similarity
    • Document classification via Naïve Bayes and Maximum Entropy
  7. Identifying Important Text Elements
    • Dimensionality reduction: PCA, SVD, and non-negative matrix factorization
    • Topic modelling and information retrieval through Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis and Advanced Topic Modeling
    • Distinguishing positive and negative sentiment
    • Item Response Theory
    • POS tagging applications: identifying people, places, and organizations
    • Advanced topic modelling: Latent Dirichlet Allocation
  9. Case studies
    • Extracting insights from unstructured user reviews
    • Sentiment classification and visualization of product review data
    • Analysing search logs to identify usage patterns
    • Text classification techniques
    • Topic modelling practices

Requirements

Familiarity with NLP principles and an understanding of how AI is applied in business contexts

 21 Hours

Testimonials (1)

Upcoming Courses

Related Categories