Context
In mid-2025, Goodnotes was redesigning a toolbar that had barely changed since 2014. The product had grown around it, and each new feature stretched it further past what it was built to support. The change was necessary, but not user-led: most users liked the existing design and had years of muscle memory around it.
The toolbar is one of the few parts of the app used by every user. If we don't get the execution right, it risks overshadowing everything else in the launch. People won't see the new features if they're frustrated by the thing they use to reach them.
I joined a month after kick-off. The team was moving quickly, but was divided: the product trio had different priorities, different research questions, and a lot of tension. I had three weeks to align the team on what still needed proving, then run the evaluative work before beta.
Outcomes and impact
The research gave the team clear priorities for a tight beta timeline. Instead of continuing to debate the redesign from different assumptions, we had a shared evidence base for deciding what needed attention before launch.
The work identified usability issues, bugs, and disruptions to key workflows. It informed 11 high-priority improvements before beta, focused on reducing the risk of a familiar part of the product feeling slower or less reliable. These included changes to tool navigation and visibility, alongside clearer iconography for core features.
One month after beta launched, toolbar UMUX was 4% higher than the previous design. The toolbar was also the most stable feature in beta.
At the start of the year, there was a lot of uncertainty around product direction and team dynamics. This research helped the team move past that.
Approach
Project timeline summary
- 1
Week 1
Build project context
Understand the team, map the disagreement, and pull together existing evidence.
Team onboarding
Stakeholder interviews
Desk research
Alignment workshop
- 2
Week 2
Evaluative testing
Run scoped studies to test first impressions and workflow friction.
Research prep
5-second tests (n=20)
Usability interviews (n=8)
- 3
Week 3
Turn findings into decisions
Synthesis, issue prioritisation, and a team-wide share-out before beta.
Synthesis and analysis
Align on priorities
Map and document insights
Team-wide share-out
Week one: Build the shared frame
The first problem was not the toolbar. It was the team's lack of a shared frame. I started by mapping the assumptions behind the disagreement across design, product, engineering, and Customer Support. In parallel, I reviewed existing research and design rationale to separate what we knew from what still needed evidence.
Design had strong views on flows and interaction clarity. Product was balancing user needs with business priorities and leadership pressure. Engineering was working towards a code freeze. I used an alignment workshop to put those tensions into a shared space, agree what we already knew, and define what the research needed to resolve.
Week two: Test what could disrupt workflows
By the end of the workshop, the team had agreed what the research needed to resolve: whether the new toolbar still felt fast, familiar, and low-friction to existing users. Existing users were the bigger concern, but we did not want to assume new users would take to the redesign without checking. I scoped two complementary studies rather than one broad sweep.
We ran unmoderated 5-second tests with new and existing users for first-impression signal, and in-person moderated usability tests with existing users to probe workflow friction. Both stayed narrow: university students only, questions tied to the workshop outcomes, and both studies run in week two.
New users responded well in the 5-second tests. With no established habits around the old toolbar, they had little to compare the redesign against. Existing users were more hesitant, and the moderated sessions showed why. Friction showed up in everyday workflows, including changes that made the transition feel disruptive when we had been aiming for a seamless one.
Week three: Turn findings into decisions
Once the findings were clear, the remaining work was activation. Research maturity was still developing on the team, so a Notion report was not going to be enough.
I handled this in two ways. First, I worked with my PM to create a prioritisation table for every insight. Each issue was mapped by severity, priority, owner, and next step, including explicit decisions not to act. The goal was to make every finding traceable to a product decision.
Second, I ran a live team-wide shareout structured around the decisions we needed to make before beta. After weeks of debate, the team needed one shared set of priorities.
Dozens joined, including most of the product org, PMs, and designers from adjacent teams. The team left with confidence in the direction, stronger buy-in on the research, and a clear plan for what needed to change before beta.

Learnings
This project changed how I think about evaluative research in teams with strong opinions. The user risk and the team risk were connected: if the team had no shared decision criteria, the findings would have been easy to debate, delay, or ignore.
Looking back, two gaps stood out.
Deeper product trio involvement
I would involve PM, design, and engineering more directly in the work itself, through targeted observation or co-synthesis, rather than keeping synthesis separate until the shareout. That would have built trust in the findings earlier and shortened the path from evidence to decision.
Defining post-launch success
We had a clear product and business goal going in, and we tracked early signals such as UMUX and beta stability. What we did less well was agreeing upfront what would count as long-term success, and how we would interpret new UX issues after ship. In the aftermath of the redesign, that made it harder to tell whether post-launch issues were regressions from the redesign, or unrelated issues surfacing later.
Since this project, I build co-synthesis into the research plan from the start so stakeholders stay close to the evidence before decisions are made. I also agree success criteria with the product trio before fieldwork, including what we will measure after launch and how we will interpret new issues. That is how I guard against the same two gaps on fast-moving teams today.