This project turned sample text into a probability model, then used that model to create output that balanced recognizable patterns with controlled randomness.
C++
Visual Studio
Markov Chains
N-Grams
Probability
StatusCompleted
FocusSoftware + probability
Core toolsC++ · Visual Studio · Markov Chains
context = "engi" next → n 0.62 next → e 0.24 next → r 0.14 engineering...
Project overview
I developed a C++ program that learns character patterns from training text using Markov chains and n-gram modeling. It calculates possible next-character probabilities and samples from those distributions to generate new text.
Technical details
How the system works.
A closer look at the architecture, implementation decisions, and validation behind this project.
01
Building the language model
The program reads training text as a sequence of characters. For each position, it treats the previous n characters as the current context, or n-gram, and records which character appears next.
Repeating that pass over the source text creates a transition model. A context that appears many times can have several possible next characters, and the number of observations for each option becomes the evidence used to calculate its probability.
02
Weighted character selection
Generation begins with a valid context from the model. The program looks up every character that followed that context during training and performs weighted sampling, so common transitions are selected more often without making the output completely deterministic.
After choosing a character, the context window moves forward by one position. The oldest character leaves the n-gram, the new character is appended, and the process repeats to produce the requested amount of text.
03
Choosing the n-gram order
The n-gram order controls how much recent history affects the next decision. A smaller order creates more possible transitions and therefore more variation, but the result can lose recognizable spelling and structure.
A larger order preserves longer fragments from the training text and usually improves local coherence. If the order becomes too large, however, the output can begin repeating the source rather than creating new combinations. Comparing several orders made this tradeoff visible in the generated results.
Lower order: more randomness and weaker local structure
Higher order: stronger local structure and less variation
The useful setting depends on the size and variety of the training text
04
Testing probabilistic behavior
Because the output is random, one successful sentence is not enough to evaluate the model. I generated multiple samples and compared how consistently each n-gram order produced recognizable patterns.
This project changed how I think about testing. For deterministic code, one input should produce one expected result. For a probabilistic program, testing also means examining distributions, repeated behavior, edge contexts, and whether randomness is weighted correctly.
Engineering approach
From idea to working system.
01
Model the context
The program uses an n-gram as its current context, allowing recent characters to influence the next generated character.
02
Build probabilities
I calculated next-character probabilities from the training text so frequently observed transitions were more likely to be selected.
03
Tune the output
I tested different n-gram orders to compare coherence and randomness, using the results to understand the model's tradeoffs.
Technical highlights
What this project demonstrates.
01
Character-level probability modeling
02
N-gram context generation
03
Weighted probability sampling
04
Comparative testing across n-gram orders
What I learned
The project gave me practical experience translating a mathematical model into a working C++ system and evaluating behavior that is probabilistic rather than fixed.