MDS Computational Linguistics

UBC’s Master of Data Science in Computational Linguistics is the credential to set you apart. Offered at the Vancouver campus, this unique degree is tailored to those with a passion for language and data. Over 10 months, the program combines foundational data science courses with advanced computational linguistics courses, like natural language processing (NLP), equipping graduates with the skills to turn language-related data into knowledge and to build AI that can interpret human language.

For information on the MDS Computational Linguistics program, please go here.

Domestic applications to MDS Computational Linguistics received after June 1, 2026 may be reviewed on a case by case basis but requires direct communication with the program. Interested applicants may contact ling.mds@ubc.ca for further information or questions.

Are you passionate about language?

Are you passionate about language and curious about data? UBC’s Master of Data Science in Computational Linguistics specialization was designed for you. An accelerated, 10-month, full-time program gets you into a career faster.

Program Benefits

Highlights Across All MDS Programs:

  • 10-month, full-time, accelerated program offers a short-term commitment for long-term gain
  • Condensed one-credit courses allow for in-depth focus on a limited set of topics at one time
  • Capstone project gives students an opportunity to apply their skills
  • Real-world data sets are integrated in all courses to provide practical experience across a range of domains

Highlights Specific To Computational Linguistics:

  • Courses are taught by a combination of arts (linguistics), computer science, and statistics faculty members giving students access to key experts within each field of study
  • Students learn fundamental data science skills, techniques, and tools with the core Master of Data Science cohort, then branch off into more specialized courses, such as natural language processing (NLP), experiencing the benefits of a large program and small program in one
  • Students learn how to solve problems involving language data (i.e., text)
  • Students learn how to use and build AI models including Large Language Models (LLMs), chatbots, Retrieval Augmented Generation (RAG) systems, and agentic workflows
  • UBC’s Vancouver campus offers students the unrivaled experience of a top 40 university, surrounded by remarkable natural beauty, at the edge of a cosmopolitan city
  • Strong connections with industry partners in public and private sectors, start-ups, and leading tech companies offer a wide range of networking/career opportunities

MDS Computational Linguistics Frequently Asked Questions - Admissions

In this video, Garrett Nicolai, Director and Assistant Professor of Teaching, for the UBC Master of Data Science (MDS) Computational Linguistics program, answers the most frequently asked question about the admissions process for the MDS Computational Linguistics program.

MDS Computational Linguistics Frequently Asked Questions - Program

In this video, Garrett Nicolai, Director and Assistant Professor of Teaching, for the UBC Master of Data Science (MDS) Computational Linguistics program, answers the most frequently asked question about the MDS Computational Linguistics program.

Curriculum*

The program structure includes 24 one-credit courses offered in four-week segments. Courses are lab-oriented and delivered in-person.

At the end of the six segments, an eight-week, six-credit capstone project is also included, allowing students to apply their newly acquired knowledge, while working alongside other students with real-life data sets. Please note that instructors are subject to change.

* subject to change at the discretion of the MDS Computational Linguistics program

Fall: September - December

Block 1 (4 weeks, 4 credits)

Programming for Data Science | DSCI 511

Program design and data manipulation with Python. Overview of data structures, iteration, flow control, and program design relevant to data exploration and analysis. When and how to exploit pre-existing libraries.
Shareable Course Page
Muhammad Abdul-Mageed

Computing Platforms for Data Science | DSCI 521

How to install, maintain, and use the data scientific software stack. The Unix shell, version control, and problem solving strategies. Literate programming documents.

Shareable Course Page
Ilya Musabirov (Section 1), Daniel Chen (Section 2)

Programming for Data Manipulation | DSCI 523

Program design and data manipulation with R. Organizing, filtering, sorting, grouping, reformatting, converting, and cleaning data to prepare it for further analysis.
Shareable Course Page
Gittu George (Section 1), Payman Nickchi (Section 2)

Descriptive Statistics and Probability for Data Science | DSCI 551

Fundamental concepts in probability including conditional, joint, and marginal distributions. Statistical view of data coming from a probability distribution.
Shareable Course Page
Payman Nickchi (Section 1), Alexi Rodriguez-Arelis (Section 2)

Block 2 (4 weeks, 4 credits)

Algorithms & Data Structures | DSCI 512

How to choose and use appropriate algorithms and data structures to help solve data science problems. Key concepts such as recursion and algorithmic complexity (e.g., efficiency, scalability).

Shareable Course Page
Muhammad Abdul-Mageed

Data Visualization I | DSCI 531

Exploratory data analysis. Design of effective static visualizations. Plotting tools in R and Python.

Shareable Course Page
Joel Östblom (Section 1 and 2)

Statistical Inference and Computation I | DSCI 552

The statistical and probabilistic foundations of inference. Large sample results. The frequentist paradigm.
Shareable Course Page
Alexi Rodriguez-Arelis (Section 1), Rodolfo Lourenzutti (Section 2)

Supervised Learning I | DSCI 571

Introduction to supervised machine learning. Basic machine learning concepts such as generalization error and overfitting. Various approaches such as K-NN, decision trees, linear classifiers.
Shareable Course Page
Varada Kolhatka (Section 1 and 2)

Block 3 (4 weeks, 4 credits)

Corpus Linguistics | COLX 521

Why do modern language models operate on tokens instead of words? What’s the difference? How do we transform large data corpora into mathematical representations that can be processed by modern AI? This course introduces the concept of textual transformation that allows modern AI to function properly.
Shareable Course Page
Garrett Nicolai

Databases & Data Retrieval | DSCI 513

How to work with data stored in relational database systems. Storage structures and schemas, data relationships, and ways to query and aggregate such data.
Shareable Course Page
Gittu George (Section 1 and 2)

Regression I | DSCI 561

Linear models for a quantitative response variable, with multiple categorical and/or quantitative predictors. Matrix formulation of linear regression. Model assessment and prediction.
Shareable Course Page
Rodolfo Lourenzutti (Section 1), Payman Nickchi (Section 2)

Feature and Model Selection | DSCI 573

How to evaluate and select features and models. Cross-validation, ROC curves, feature engineering, and regularization.
Shareable Course Page
Prajeet Bajpai (Section 1), Elham E Khoda (Section 2)

Winter: January - April

Block 4 (4 weeks, 4 credits)

Parsing for Computational Linguistics | COLX 535

Learn the linguistic theory behind grammatical structure - the techniques that make large language models possible. This course introduces students to the analysis of linguistic output, interpreting the grammatical structure for deeper meaning that can help identify shortcomings in AI output.
Shareable Course Page
Garrett Nicolai

Computational Semantics | COLX 561

An introduction to Large language models (LLMs). How do we get AI models that can understand and produce human language? This course introduces the idea of creating semantic vector spaces through language generation, as well as the architectures used in LLMs.

Shareable Course Page
Isabel Papadimitriou

Unsupervised Learning | DSCI 563

How can we learn the answers to problems that don’t have labels? This course introduces students to a sequence of unsupervised methods for clustering, dimensionality reduction and vector visualization, and teaches students to find structure in messy data.
Shareable Course Page
Garrett Nicolai

Supervised Learning II | DSCI 572

How do modern neural networks learn? This course introduces the mathematical foundations of deep learning and optimization, Topics include optimization methods from gradient descent to Adam and Muon optimizers, architectures ranging from fully connected networks to transformers, and various methods for language modeling, with a focus on mathematical derivation and code implementation.

Shareable Course Page
Jian Zhu

Block 5 (4 weeks + 1 week break, 4 credits)

Advanced Corpus Linguistics | COLX 523

We have the data, so what can we do with it? From annotation to deployment, this course introduces students to data stewardship and ownership. Students create and deploy a corpus and dashboard that is informative and accessible, all while taking ownership through an AGILE framework.

Shareable Course Page
Garrett Nicolai

Computational Morphology | COLX 525

What happens when the word is not small enough, and our systems need to pick them apart further? Learn how linguistics has defined “meaning bearing units”, and how we adapt that knowledge to gain insight about how to interpret the output of AI in several modalities: text and speech.
Shareable Course Page
Garrett Nicolai

Machine Translation | COLX 531

How are large language models actually built and deployed? This course covers the major stages of modern LLM engineering, including how to scale up in pre-training, how methods like RLHF and RLVR are used in post-training, and how to design LLM inference engines for real-world deployment.
Shareable Course Page
Jian Zhu

Sentiment Analysis | COLX 565

A research seminar-style class, reading recent papers in the field and discussing their implications. Main focuses are interpretability (methods for understanding what’s happening in the black box of LLMs) and the effects of LLMs on human psychology and society.

Shareable Course Page
Isabel Papadimitriou

Block 6 (4 weeks, 4 credits)

Advanced Computational Semantics | COLX 563

Topics in LLMs. This course connects the fundamentals learned in COLX 561 and DSCI 531 to some of the concerns of building and using LLMs today. Topics include scaling laws, GPUs, embedding models, and architectural choices.

Shareable Course Page
Isabel Papadimitriou

Natural Language Processing for Low-Resource Languages | COLX 581

We’re in the Big Data revolution, but not all languages have equal amounts of data. Learn how to adapt large language models to smaller data, including efficient methods of selecting from limited examples, bootstrapping and ensembling, LoRA and PEFT, and in context and few-shot learning.

Shareable Course Page
Garrett Nicolai

Trends in Computational Linguistics | COLX 585

The latest developments in NLP and AI. Topics change each year to follow major new directions in the field. This iteration focuses on recent innovations in neural architectures, LLM agents, reasoning and tool use, and multimodal understanding and generation.
Shareable Course Page
Jian Zhu

Privacy, Ethics & Security | DSCI 541

How do we establish fairness while also preserving security? What is AI doing with your data, and how should we approach questions concerning digital colonialism?
Shareable Course Page
Garrett Nicolai

Spring: May - June

Capstone Project (8-10 Weeks, 6 credits)

Capstone Project | COLX 595

A mentored group project based on real data and questions from a partner within or outside the university. Students will formulate questions and design and execute a suitable analysis plan. The group will work collaboratively to produce a project report, presentation, and possibly other products, such as a web application.


Shareable Course Page
MDS Computational Linguistics Staff

Meet Amy

Even though Amy found the MDS Computational Linguistics program an intensive and accelerated one, it actually better fit her needs. Amy felt the most important thing they learned is to solve problems and once you are able to see a clear picture of the data, you are able to feel a sense of achievement.