02 / Software + probabilityNov - Dec 2025

Markov Chain Text Generator

This project turned sample text into a probability model, then used that model to create output that balanced recognizable patterns with controlled randomness.

  • C++
  • Visual Studio
  • Markov Chains
  • N-Grams
  • Probability
StatusCompleted
FocusSoftware + probability
Core toolsC++ · Visual Studio · Markov Chains

Project overview

I developed a C++ program that learns character patterns from training text using Markov chains and n-gram modeling. It calculates possible next-character probabilities and samples from those distributions to generate new text.

Technical details

How the system works.

A closer look at the architecture, implementation decisions, and validation behind this project.

01

Building the language model

The program reads training text as a sequence of characters. For each position, it treats the previous n characters as the current context, or n-gram, and records which character appears next.

Repeating that pass over the source text creates a transition model. A context that appears many times can have several possible next characters, and the number of observations for each option becomes the evidence used to calculate its probability.

02

Weighted character selection

Generation begins with a valid context from the model. The program looks up every character that followed that context during training and performs weighted sampling, so common transitions are selected more often without making the output completely deterministic.

After choosing a character, the context window moves forward by one position. The oldest character leaves the n-gram, the new character is appended, and the process repeats to produce the requested amount of text.

03

Choosing the n-gram order

The n-gram order controls how much recent history affects the next decision. A smaller order creates more possible transitions and therefore more variation, but the result can lose recognizable spelling and structure.

A larger order preserves longer fragments from the training text and usually improves local coherence. If the order becomes too large, however, the output can begin repeating the source rather than creating new combinations. Comparing several orders made this tradeoff visible in the generated results.

  • Lower order: more randomness and weaker local structure
  • Higher order: stronger local structure and less variation
  • The useful setting depends on the size and variety of the training text
04

Testing probabilistic behavior

Because the output is random, one successful sentence is not enough to evaluate the model. I generated multiple samples and compared how consistently each n-gram order produced recognizable patterns.

This project changed how I think about testing. For deterministic code, one input should produce one expected result. For a probabilistic program, testing also means examining distributions, repeated behavior, edge contexts, and whether randomness is weighted correctly.

Engineering approach

From idea to working system.

01

Model the context

The program uses an n-gram as its current context, allowing recent characters to influence the next generated character.

02

Build probabilities

I calculated next-character probabilities from the training text so frequently observed transitions were more likely to be selected.

03

Tune the output

I tested different n-gram orders to compare coherence and randomness, using the results to understand the model's tradeoffs.

Technical highlights

What this project demonstrates.

01

Character-level probability modeling

02

N-gram context generation

03

Weighted probability sampling

04

Comparative testing across n-gram orders

What I learned

The project gave me practical experience translating a mathematical model into a working C++ system and evaluating behavior that is probabilistic rather than fixed.
Next projectMicrocontroller Tic-Tac-Toe