The human layer: Appen’s critical role in the AI ecosystem
The Human Layer: Appen’s Critical Role in the AI Ecosystem
Appen CEO Ryan Kolln explains how a company that began collecting Australian accents for a horse racing betting line became a key data supplier to the world’s leading AI labs in both the US and China.
Appen’s Vantage Point
- Appen (ASX: APX) supplies the training data used to build and refine leading AI models, working across both the US and China ecosystems.
- The company is in its 30th year, has a crowd of more than a million contributors in 170 countries, and adds about 1,000 sign-ups a day organically.
- Ryan became CEO in early 2024, after the loss of a contract worth about a third of revenue, a 41% one-day share price fall and the previous CEO’s resignation.
- Since then he has rebuilt the leadership team, bringing in more than 20 senior hires from customers and competitors, and made quality the company’s “North Star.”
- Appen China has grown from around US$5 million a year to a run rate above US$175 million, built as a fully standalone operation.
- Generative AI work is starting to offset the decline in Appen’s traditional business, and Q2 came in ahead of Q1.
- Internal AI automation has enabled a $12 million cost-out.
- Ryan sees synthetic data as a useful tool but warns about “model collapse” if it is over-relied on. He is more interested in what the AI leaders are building than in who tops the benchmarks.
The Hidden Supply Chain Behind Every AI Model
Every AI model you use rests on a supply chain most people never see. For a model to answer well, someone has to decide what a good answer looks like, in the right language, cultural register and subject area. Appen has done that work since 1996.
Ryan describes the AI stack in three parts: the model architecture, the compute, and the data layer that shows the model how to behave. The labs employ world-class experts in maths, algorithms and running compute at scale, but they often can’t reach the people with the specific language, domain or reasoning skills they need.
Our customers, effectively the researchers who we work with, they are experts in the math and they’re experts in the algorithms and how to get the compute working at really large scale. What we bring is the expertise in how do we get a distributed network of humans to deliver really high quality data that is within a consistency barrier.
— Ryan Kolln
Ryan traces the idea back to the original ChatGPT. GPT-2 was smart but clunky, and the breakthrough was reinforcement learning with human feedback: large volumes of real human examples showing the model what a good structure and a good argument look like, so it could sound and act more like us.
From Horse Racing Bets to Humanoid Robots
Appen’s first project came in 1996. Its founder, linguistics professor Julie Vonwiller, was approached after a US speech recognition system that let callers place bets on horse races failed to cope with Australian accents. Appen collected Australian-accented English to fix it.
Ryan is quick to say today’s work is far removed from the “click the buses” puzzles people picture.
The type of work that we do is actually really, really sophisticated. Some of the tasks that we work on might take 8 hours to complete.
— Ryan Kolln
His examples show the range:
- Finance: Experts read five years of 10-Ks and 10-Qs and write questions and answers that teach models how to analyse strategy and financial outcomes over a long horizon.
- Coding: Cybersecurity consultants review model-generated code for vulnerabilities, and the findings are fed back into training.
- Medicine: Medical experts check whether the dosage instructions a model suggests match the pharmaceutical spec sheets.
- Speech: Models still break down when several people talk at once. Simon’s quip that “humans love talking over each other” sums up the problem.
- Humanoid robotics: Appen collects and annotates data for robots learning household chores like folding laundry, and for commercial settings such as kitchens and workshops. A mechanic wears a head-mounted camera while completing the task.
A Million Contributors, Matched by Machine
The obvious question is how you keep a million people on your books without blowing out costs. Ryan says the answer is heavy automation. Contributors sign up, describe their background and interests, and are matched to tasks with a pay rate attached. Pay varies by location and the capability required.
Growth is almost entirely organic. There are Reddit forums and a whole community of part-time workers, many of them professionals working weekends, while some treat it as their main income.
A lot of people are professionals, and they’re doing it on the weekend. Some people do it as a full-time, you know, that is their primary source of income.
— Ryan Kolln
On the revenue side, Appen sells units of data. A customer asks for, say, 10,000 units, Appen sets up the project, finds the people and delivers, and gets paid for what it provides. Some projects have a fixed end date, while others have run for ten years. Ryan also pointed listeners to crowdgen.com to sign up.
Walking Into the Storm: The Hardest Calls as CEO
In January 2024, Appen lost a contract worth about a third of its revenue. The stock fell 41% in a day, and within two weeks the CEO had resigned. Ryan, then COO, took over the same day. He was excited to lead but says the circumstances “weren’t entirely optimal.”
His focus from day one was getting back to fundamentals: doing right by customers and rebuilding a culture around high-quality data delivered quickly. The hardest part was people. The market had changed, and some talented long-serving colleagues didn’t have the capabilities the next phase demanded.
There have been people in the company who are very talented and were very well-suited for the time that they were in Appen, but perhaps they didn’t have the capabilities to evolve with the market to the future state.
— Ryan Kolln
The replacements came from an unusual place. Ryan hired more than 20 senior people from both customers and competitors, a deliberate talent strategy built around the way frontier researchers work. They are constantly experimenting, and their data needs are often things no one has done before.
You get in early, you build that relationship, you build that trust, and that’s what expands into the bigger projects and ultimately the better financial outcomes.
— Ryan Kolln
He says the shift was less about sales or culture than about assembling people who can deliver quality on complex, domain-specific tasks. “The rest looks after itself when it’s of such good quality.”
The Turnaround Signal: Generative AI Up, Legacy Work Down
Appen Global has seen its traditional work decline as AI automates some of those functions, while large language model work has grown. The signal Ryan watches is the quarter where growth in new areas outweighs the decline in old ones. Q2 was up on Q1, which he calls a really good sign.
At the same time, the company has cut costs, and much of that comes from using AI on itself in two main ways:
- Quality control: Multiple AI checkers, sometimes 10 to 20 specially trained models, run across contributor data.
- Software development: Appen builds its own contributor portal and annotation tools rather than buying from third parties, with a much leaner engineering team.
This underpins the recently announced $12 million cost-out. Ryan says the goal is agility as much as dollars, since a smaller team backed by powerful AI models can “flex super fast” to the needs of researchers.
Synthetic Data: Threat or Tool?
Synthetic data is often raised as the bear case for a business like Appen. Ryan doesn’t see it as a threat. It has been part of the market for a long time and Appen uses it too. A mix of synthetic and human data is now the standard approach.
He does flag the risk of overdoing it.
If you over-rely on synthetic data, we’re effectively getting a model to train a model. You can get a concept called model collapse, where it gets in a reinforcement loop that doesn’t head in the right direction.
— Ryan Kolln
His broader point is that labs use many techniques, including synthetic data and scraping the internet. As models improve they need more techniques, and human data is one of them. “It’s a rising tide across all areas.”
The China Bet
Appen started in China in 2018, when the industry was nascent. It seeded the team with expertise from its US customer work, but made a deliberate decision to run a fully standalone operation, with its own team, technology stack, finance and HR, in order to move at “China speed.” Ryan says China was around US$5 million for the year not long ago. It is now running past a US$175 million annualised run rate.
Appen China serves the leading Chinese labs, both for domestic applications and for large Chinese technology companies operating globally, which tap into Appen’s international network. Ryan says the early obstacle wasn’t being a Western-listed company, it was proving they could deliver the data. He describes the country with real enthusiasm.
Every time I go there, I feel like I’ve stepped into the future.
— Ryan Kolln
Appen is now one of the bigger players in China and has passed some established competitors. Ryan credits three decisions:
- Pay well and retain talent, so expertise didn’t leak to competitors and customers.
- Empower local leadership, with corporate controls in place but no centralised strategy calls.
- Accept duplication, including a separate technology stack that looks inefficient on paper but has proved a differentiator for responsiveness and data privacy.
On the size of the prize, he says the AI market will be “much bigger than it is now,” with 100x more likely than 3x, and the open question being timing.
Safety, the Race and the Compute Squeeze
Simon asked about this week’s doomsday commentary from former Anthropic staff, who argue that the safeguards aren’t in place for the moment models begin improving themselves. Ryan’s view is that what matters is what models are optimising for.
We just need to make sure that we’re optimizing for the right thing, which can be difficult when you’ve got a lot of investors who want a really strong commercial outcome versus what’s going to be the right thing to be doing for humanity.
— Ryan Kolln
On who wins the race, he separates consumer from enterprise AI. In consumer, he thinks distribution will decide it, favouring players like Google, Meta, X/SpaceX and possibly OpenAI. Enterprise is where companies like Anthropic sit, and there he asks whether selling tokens will remain the model or whether the labs will move up into applications. He floats a “company factory” future, such as a fully digital neobank run entirely by agents with no humans involved. He says he’s less interested in who tops which benchmark than in where all this ends up.
On compute, he acknowledges a real constraint in chips and data centres, but expects supply to catch up with demand eventually. He also points to efficiency breakthroughs, citing a partner called Subquadratic whose model uses about 5% of the compute for the same outcome, effectively a 20x boost to existing data centre capacity. In his view, the biggest step changes will come from new architectures.
What He’s Learned, and What Success Looks Like
Asked whether he’s a better CEO than two and a half years ago, Ryan says the big lesson was how much change an organisation can absorb. He was initially reluctant to push too hard, but has learned that with the right people, “super passionate about the industry, super smart, really agile,” the velocity of change can be extremely high.
Looking ahead 12 months, he says the numbers are an outcome. What he watches is the depth of relationships with researchers, becoming a more important part of their ecosystem, and making Appen lean and agile enough to move as fast as they do.