Back to Insights Product

How We Trained on 200,000 Professional Emails

Nadia Ferretti 7 min read
How We Trained on 200,000 Professional Emails

When we decided to build a grammar and style tool specifically for professional writers, we had to make a dataset decision that would shape everything downstream. The standard corpora used to train language models and grammar tools are dominated by academic writing, news text, and literary fiction. These are the sources that are clean, well-edited, and easily accessible. They are also, from the perspective of someone who writes workplace email and documents, the wrong kind of English.

Professional email is not academic writing. It is not even close. The sentence structures are shorter, the register shifts rapidly depending on context and relationship, the tolerance for informality is much higher, and the error patterns that matter are different. A model trained on academic prose will flag plenty of constructions that are perfectly normal in professional email and miss plenty of constructions that create real friction in that context.

Why we needed a different dataset

The specific patterns we care about at Linguix are patterns that affect how professional writers are perceived by the people they write to: colleagues, clients, managers, candidates, partners. Those people are reading workplace email. They are not reading journal articles.

This matters for several concrete reasons. Register is the most important one. The appropriate level of formality in workplace email in 2025 is significantly different from what it was in 1995, and it differs substantially from the register of academic or formal writing. A model that associates "professional writing" with "formal academic writing" will produce suggestions that push writers toward an outdated register, which is not helpful and can be actively counterproductive.

Error patterns are the second reason. The grammatical errors that appear in professional email, particularly email written under time pressure, differ from the errors that appear in academic writing. Subject-verb agreement errors in complex sentences, preposition errors in phrasal verb constructions, and article errors around specialized terminology all pattern differently across these genres. Training on academic text to correct professional email errors is like training on French cuisine to cook Brazilian food: related, but not the same.

What we collected and how

The core of our training dataset is a collection of approximately 200,000 professional email exchanges from publicly available datasets and licensed sources, focused specifically on workplace communication in English-language business contexts. We supplemented this with a smaller but carefully annotated set of corrected writing samples from early Linguix users who opted into the feedback program.

We did not use synthetic data for the core grammar modeling layer. This was a deliberate decision. Synthetic data, at this stage, does not yet reproduce the distributional properties of real professional writing well enough to be useful for training nuanced style suggestions. The error patterns in real writing are messier and less predictable than in synthetic data, and that messiness is exactly what we need to model.

What we did use synthetic data for was data augmentation around specific rare error categories. Some errors, like certain preposition confusion patterns, appear frequently enough in real email but not uniformly enough across different industries and communication contexts to give us adequate coverage from real data alone. For these cases, we generated targeted examples to fill out the distribution, validated against real examples to check that the synthetic cases were realistic.

The cleaning problem

Professional email data comes with significant noise. Forwarded chains, quoted text, email signatures, automated system messages, boilerplate legal footers. All of this has to be identified and stripped before training, because otherwise the model learns from the wrong text. We spent more engineering time on cleaning than on any other single part of the data pipeline.

A less obvious cleaning challenge is the identification of intentional informal writing versus unintentional errors. Professional email deliberately uses informal constructions that look like errors but are not: incomplete sentences as deliberate fragments, comma splices as stylistic choices, sentence-starting conjunctions for emphasis. A training dataset that marks all of these as errors will produce a model that flags legitimate professional style choices. Differentiating intentional informal from unintentional error requires context and annotation that raw scraped email does not provide.

We handled this partly through annotator training, asking annotators to evaluate constructions in context rather than in isolation, and partly through a two-stage pipeline that first classifies writing context (casual exchange vs. formal request vs. client-facing communication) before applying correction labels. This does not solve the problem completely, but it substantially reduces the rate of over-flagging informal-but-intentional writing.

What the email focus means for coverage

We want to be clear about what this means for Linguix's strengths and limitations. Our models perform well on the kinds of writing that professional email represents: direct workplace communication in English, covering requests, updates, feedback, scheduling, and relationship management. They perform less well on highly specialized technical writing, literary writing, and formal academic writing, because those are not the contexts we trained on.

We also recognize that "professional email" is not a homogeneous category. A cold sales email has different norms than an internal team update. An email to a client in a regulated industry has different register requirements than a message to a startup founder. The 200,000 email dataset captures a range, but it is not infinite, and there are professional communication contexts where our suggestions will be less calibrated than we would like.

What we learned from early users

When we rolled out the initial Linguix beta, one of the most useful feedback patterns we saw was users identifying cases where the suggestions were wrong in instructive ways. Not "this is incorrect" but "this would make my email sound weird to my team." Those cases, collected and analyzed, told us more about the gaps in our training data than any internal evaluation did.

The most common pattern: suggestions that were correct for formal business communication but wrong for the user's actual communication context. An internal team update getting flagged for being too informal. A casual project check-in getting suggestions that would push it toward a client-facing register. Both of these are symptoms of a single underlying issue: insufficient granularity in how we model communication context. They are also the things we focused on in the iteration following the initial beta.

Building a good grammar and style model for professional writing turns out to be mostly a data problem. Get the right data, clean it properly, label it carefully, and you have most of what you need. Everything else is model architecture and fine-tuning. The data step is where the real decisions are made.

Try Linguix in your browser

Real-time corrections, right where you write. Free to install.

Add to Chrome - Free More Insights