Text, Tokenization, and Embeddings
introduce text workflows without jumping straight to large language models.
Units
- Unit 09.00: 09.00 Text as data and why tokenization matters
- Unit 09.01: 09.01 Vocabulary, sequence length, padding, and truncation
- Unit 09.02: 09.02 TextVectorization and simple baselines
- Unit 09.03: 09.03 Embedding layers and learned representations
- Unit 09.04: 09.04 Comparing text representations honestly
- Unit 09.05: 09.05 Project step: text classification or embedding report
