Recluse Studio
Field note / Authored record
StudioBlogSupport
← Field notes

Enthusiasm Is Not a Stable Measure of AI Adoption

During an eight-week Copilot pilot, skeptics often became more positive while champions became less enthusiastic and employees narrowed the tasks they trusted.

The same worker moves from champion enthusiasm through hands-on testing to calm judgment, keeping a useful summary and discarding a weak chart.
Post-specific field image / portrait

Scope note: This essay covers changing employee perceptions during an eight-week Microsoft 365 Copilot pilot at one state transportation agency. It does not measure long-term use or objective productivity.

The skeptic and the champion may both be wrong before they use the tool. Experience moves them toward the middle for different reasons.

A new study followed employees at the North Carolina Department of Transportation through an eight-week Microsoft 365 Copilot pilot. The researchers matched surveys from before and after use, leaving 124 responses after quality checks. They tracked usefulness, ease, intent to keep using the tool, and trust.

External record / arxiv.orgPersona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of TransportationGenerative AI tools are increasingly being piloted in public agencies, but limited evidence explains how employee acceptance changes after hands-on use. This study examines Microsoft 365 Copilot adoption during an eight-week pilot at a sta…

Average usefulness fell. Most other measures barely moved. Under those averages, individual employees changed position sharply.

Hands-on use removed some optimism and some fear

Before the pilot, the researchers grouped employees into three profiles: Skeptics, Cautiously Positive users, and Champions.

After eight weeks, 40 percent of the Skeptics moved into the Cautiously Positive group. At the same time, 68 percent of the Champions became less enthusiastic. Most moved one step toward caution, while a smaller group moved all the way to Skeptic.

The workforce did not simply become more positive or more negative. It became more informed. Initial hopes met actual limits. Initial doubts met actual uses.

That pattern is easy to miss in an average. The overall number of people in each group changed only modestly because movement happened in both directions. A launch survey could have labeled people and left the organization with a neat, false map.

Adoption profiles are not personality types. They are temporary positions shaped by experience.

Employees kept the uses that survived contact

Communication and summarization remained relatively steady during the pilot. Presentation work, data analysis, and chart generation declined after employees tried them.

This is not necessarily failure. Dropping a weak use can be evidence of better judgment. An employee who stops asking Copilot for charts after checking the results may be calibrating trust correctly.

The concern pattern moved too. Worries about accuracy and data privacy declined, while concerns about jobs and skills increased. As employees became more familiar with immediate system behavior, their attention shifted toward what the tool might mean for their work over time.

One training session cannot cover that movement. Early support may need to explain safe use and verification. Later support may need to address workflow changes, role expectations, and skill development.

A falling score can contain progress

Perceived usefulness dropped from 3.85 to 3.62 on a five-point scale, a statistically significant change. Ease of use, intention, and trust changed only slightly.

It would be tempting to call the decline disappointing. I think it is more useful to call it a corrected estimate.

Early excitement is cheap because it has not yet paid the cost of checking output, repairing a weak draft, or discovering that a polished presentation contains poor analysis. A lower but more accurate view of usefulness gives the organization a better foundation for deciding where the tool belongs.

The goal is not maximum enthusiasm. It is calibrated use.

Eight weeks is still early

The study covers one public agency and relies on self-reported ratings. It does not connect those ratings to Copilot activity, task results, or saved time. Eight weeks cannot show whether the new views persist. Some movement paths involved very few people, and the open-text analysis used keyword matching that may have missed nuance.

Those limits point toward a better adoption dashboard. Track the same people over time. Connect perceptions to actual task use and review behavior. Look at which uses remain, which disappear, and which require correction. Do not assume that a high score means good judgment or that a lower score means rejection.

The most credible user may be neither the champion nor the skeptic. It may be the person who has learned exactly where the tool stops helping.