Enterprise productivity report

Measure Real Productivity Across Your Entire Enterprise

Get a Demo

How to Calculate an Employee Productivity Score with Microsoft 365 Data

Microsoft's Productivity Score is now Adoption Score. See what it measures, then build your own 0 to 100 productivity score from Microsoft 365 data.

Short answer: An employee productivity score from Microsoft 365 data is a weighted 0 to 100 number built from work-pattern metrics such as meeting load, focus time, and workday length. Collect calendar, email, and Teams metadata, convert each metric to a 0 to 100 scale, weight the metrics, and report by team. Microsoft's own Adoption Score is different. It measures how an organization uses Microsoft 365 tools, not how productive people are.

Most guides on this topic stop at a list of metrics and a sample formula. The hard part is that a score is only as good as its definitions and cut-offs, and those are rarely stated. Two teams can run the same formula on the same Microsoft 365 tenant and get different scores because one counts overlapping meetings twice or measures focus time in longer blocks.

This guide is for IT and HR analytics teams. It shows what Microsoft's own score measures, the definitions that quietly change a score, the thresholds Worklytics has published for burnout and focus, and a scoring query built on them. The query was run on a synthetic test dataset, and the page says so wherever it uses those results. Worklytics sells workplace analytics software, so the guide also points out where a tool helps and where you can do the work yourself.

Five-step diagram: collect metadata, clean and group it, convert each metric to a 0 to 100 score, weight and add the scores, then track the trend.
The five steps covered in this guide.

What Microsoft's Productivity Score Measures (and What It Does Not)

Many people who search for a Microsoft productivity score want the report inside the Microsoft 365 admin center. Microsoft launched it as Productivity Score in 2020 and replaced it with Adoption Score in 2022. Its purpose is to show how an organization uses Microsoft 365, not to rate individual employees. The name change followed a backlash: the 2020 version could show user names, and on December 1, 2020, Microsoft announced it would remove them and show only organization-level data. Anyone building their own score should take the same lesson about employee trust.

The report changed again in 2026. Starting January 22, 2026, Microsoft retired the Technology experiences part (network connectivity, Microsoft 365 Apps health, and endpoint analytics). The score now adds up five categories worth 100 points each, plus AI adoption if Copilot licenses are enabled, for a maximum of 600 points. It reflects the last 28 days and is calculated at the organization level, never for individuals. Older articles that list technology categories or other point totals describe earlier versions. To track Copilot use in more detail, see how to track Copilot utilization.

Adoption ScoreViva InsightsCustom score (this guide)
The question it answersHow well does our organization use Microsoft 365?How do people and teams spend their time: meetings, focus, after-hours work?How do the work-pattern metrics we care about add up to one number we can track by team?
Level of detailOrganization, with group filters based on Microsoft Entra IDPersonal insights for employees, team insights for managers, organization insights for leaders, and analyst queriesTeam or department
Who sees itAdmin roles such as Global Administrator or Reports ReaderEmployees see their own insights. Analysts need the Insights Analyst roleWhoever you choose. Keep it to managers and leaders, at team level

Adoption Score answers a technology question. A custom score answers a different one: do our teams work in a way that leaves room for focused work, avoids overload, and stays sustainable? It shows how time is used, not the quality or value of the work, so treat it as a signal that starts a conversation, not as a grade for a person.

What the Data Looks Like Before You Score It

The bigger risk sits upstream of the formula. The metrics you feed into it depend on definitions that differ between tools. This section uses Worklytics's public data documentation to show the choices that change a score.

Worklytics exports weekly aggregates as one row per person, week, and metric, with four fields: employeeId, week, key, and value. Groups such as team, level, and employment type sit in a second table. That long format is why a scoring query is a set of filtered sums grouped by team, and why the minimum group size can be enforced in the query itself.

QuestionHow Worklytics answers itWhy it matters for a score
Who counts as active?A person is active in a week if they took any action in digital tools other than a calendar attendance or an email auto-response.People on leave, or who only accept invites, drag team averages toward zero. Worklytics's own sample SQL also requires a weekday span of at least 4 hours and filters to full-time staff.
What counts as a meeting?A calendar event with at least 2 human attendees who have not declined, under 12 hours long. Tentative and unanswered invites count as attending, and meeting rooms do not count as attendees.Declined invites and one-person calendar blocks do not inflate meeting load, and a meeting people attend without accepting is still counted.
How is meeting time counted?Four layers: meeting count, direct time (every meeting's hours added up), exclusive meeting time (overlaps counted once), and exclusive time (also subtracts time spent in other tools during the meeting).Two overlapping one-hour meetings are 2 hours of direct time but 1 hour of exclusive meeting time. Dividing direct time by work hours overstates the load of overbooked calendars. Worklytics recommends meeting count or exclusive meeting time in most cases.
What is after-hours work?By default Monday to Friday, 9:00 to 17:00, adjusted to each person's time zone from calendar data. Work hours from an HRIS file, or hours people set in their own calendar, replace the default.Until real work hours are loaded, shift workers, part-timers, and people in other time zones look like they work after hours.
What is focus time?Several variants: uninterrupted blocks of 30 minutes or more, 1 hour or more, or 2 hours or more, each free of meetings, Slack, and email.Block length changes the result by role. See Step 4.
What is a strong collaborator?Someone a person spent at least 2 hours collaborating with in a week.It is the basis for the isolation signal in the next section.

Viva Insights has its own definitions. For example, uninterrupted hours are blocks of one hour or longer, while the main Worklytics focus time metric uses blocks of two hours or longer. Neither is wrong, but the numbers are not interchangeable. Choose one source of truth for the score, write the definitions down, and do not mix tools inside one score.

The Thresholds Behind the Score

A score needs cut-offs: how many focus hours is good, how long a workday is too long. Instead of inventing them, this guide uses thresholds Worklytics has already published in two posts, a May 2022 post on burnout indicators and the October 2023 benchmarks post.

IndicatorPublished thresholdHow the score uses it
Workday lengthMore than 9 hours a day is a burnout indicatorA person-week is flagged when the weekday span is over 9 hours
Weekend workMore than 2 hours on weekends is a burnout indicatorA person-week is also flagged when weekend work is over 2 hours
Focus time3.5 hours a day or more goes with people reporting they are more productive (2023). The 2022 post used 3 hours3.5 hours scores 100. The 0.5 hour floor for a score of 0 is this guide's own choice
IsolationFewer than 5 strong collaborators in a week is a proxy for isolation, and fewer than 3 is a red flagConnection is the share of people with 5 or more strong collaborators
Organizational overhead12 or more strong collaborators tends to go with higher burnout risk and attritionNot scored. Watch it next to the score
Meeting loadNo published threshold. The benchmarks post defines the normal range as the middle 50% of peopleScored against the middle 50% of meeting load across everyone in your data

Read these as starting points. The posts do not publish sample sizes or how the thresholds were derived, and the focus threshold moved from 3 hours in 2022 to 3.5 hours in 2023. The benchmarks post also warns that a normal range is not always a good range: a company can sit at the median for focus time and still have too little of it. Check the thresholds against your own survey data before you rely on them, and write down which ones you used.

Step 1: Get the Data and Set Privacy Rules First

There are three common routes to Microsoft 365 collaboration data:

  1. Viva Insights analyst tools. A person with the Insights Analyst role runs a query or a Power BI template such as Ways of working, which reports collaboration, meeting, focus, and coaching patterns by week.
  2. Export to Microsoft Fabric. Microsoft recommends the Viva Insights connector for Microsoft Fabric to move query results into storage such as a lakehouse or Azure SQL.
  3. Microsoft Graph directly. Microsoft Graph exposes calendar events, mail metadata, and Teams call and chat records, and you calculate the metrics yourself. Request metadata only, never message bodies. Worklytics reads these records through Microsoft Graph, with an optional data loss prevention proxy that runs in the customer's own environment and removes or pseudonymizes fields before data is sent.

Set privacy rules before you extract anything. Always aggregate data so that no group is small enough to identify one person. Microsoft requires Viva Insights administrators to set the minimum group size to at least five, and many organizations choose a higher number, such as 10, for scores that leaders compare across teams. Pick one threshold, apply it everywhere, and write it down. This approach supports compliance with GDPR, CCPA, and other data protection laws, but the rules depend on your country and your data, so involve your legal team. The EU AI Act also lists AI systems used to monitor and evaluate the performance and behavior of workers among its high-risk uses, so ask legal whether your scoring tool falls under it.

  • Show team scores to managers, never individual scores, and use weekly or monthly aggregates instead of daily data for any person.
  • Decide what happens to small teams before you score. With a minimum group size of five, a four-person team never appears in the results. Either roll small teams into a parent group or accept that part of the organization has no score.
  • Suppress sensitive keywords in meeting titles and email subject lines. Viva Insights has keyword and domain suppression settings for this, and an end-user opt-out that removes a person's behavioral metrics from row-level outputs.
  • Tell employees what is measured and why, and give them a way to ask what data is held about them.

Step 2: Choose the Metrics and Map Them to Your Data

True productivity covers efficiency, effectiveness, and sustainability. Microsoft 365 data covers some of that well and some of it not at all. The table shows which metrics this guide scores, where each one comes from in Viva Insights and in the Worklytics data dictionary, and what needs a second data source. For a wider list of metrics, see employee productivity metrics and KPIs.

MetricUsed forViva Insights sourceWorklytics key
Focus time (blocks of 2 hours or more)Focus scoreUninterrupted hours (blocks of one hour or longer)worklytics:hours:in:focus:blocks:v3_5:flow
Meeting hours per weekMeeting load scoreMeeting hourscalendar:events:hours:meetings
Strong collaboratorsConnection scoreNot in the metric list we reviewedcollaborators:strong_count_distinct
Workday spanBalance score, and the divisor for meeting loadNot in the metric list we reviewedworklytics:weekdays:avg:timespan:hours
Weekend workBalance scoreAfter-hours collaboration hoursworklytics:weekends:total:time:worked:hours
Large meetings, fragmented timeNot scored. Useful for diagnosisLarge and long meeting hours (nine or more invitees), interrupted hourscalendar:v1:meetings:attendees:10-or-more:hours and worklytics:hours:fragmented:v3_5:focus
Task completion, meeting outcomesNot in Microsoft 365 dataNot availableTake from your project tool or a short survey

Step 3: Clean the Data and Set a Baseline

  • Keep active people only: Worklytics's own sample SQL keeps active, full-time employees whose weekday span is at least 4 hours, and the query in this guide does the same.
  • Handle weekend schedules: if some groups work weekends by contract, leave them out of the balance component, because the weekend threshold would flag them every week.
  • Leave out non-knowledge workers and holiday weeks: Microsoft's Ways of working report can exclude people who spend under five hours a week in meetings, email, and Teams calls and chats, and weeks that look like holidays or leave.
  • Fix organization data: Viva Insights warns that low-quality or missing organizational data, such as manager or department fields, can cause errors. Fix those fields before you group by team.
  • Set a baseline: capture four or more weeks of data before you start scoring, so you have something to compare against.

Step 4: Score Focus Time

Focus time gets the largest weight in this guide's sample model, and there is published evidence behind that choice. Worklytics's research found that knowledge workers with at least 3.5 hours of daily focus time tend to report being more productive than those with less. A chart on Worklytics's Microsoft Teams analytics page shows the same pattern for engineers: those who strongly agreed that their team can work at optimal speed had a median of 4.1 hours of uninterrupted focus time a day, against 3.2 for neutral and 1.9 for those who strongly disagreed. The chart does not state its sample size or dates, so read it as illustrative evidence, not as a benchmark.

Dot plot of hours of uninterrupted focus time per day by survey answer to the statement that the team can work at optimal speed: median 1.9 hours for strongly disagree, 3.2 for neutral, and 4.1 for strongly agree.
Engineers who report higher team velocity have more uninterrupted focus time. Source: Worklytics, Microsoft Teams data analytics page.

One block length does not fit every job, and the Worklytics methodology says so. It defines separate variants for different kinds of work:

Block lengthMetric keyBest for
30 minutes or moreworklytics:hours:in:focus:blocks:v3_5:prepSales and customer success, where meetings are core work but people still need time to prepare
1 hour or moreworklytics:hours:in:focus:blocks:v3_5:focusPeople managers and senior leaders, who need to move larger projects despite heavy meeting loads
2 hours or moreworklytics:hours:in:focus:blocks:v3_5:flowDeep work, and the default for most teams. A calendar-only variant counts meetings alone

Scoring a sales team and an engineering team on the same block length makes the sales team look worse for doing its job, so pick the variant by role and compare teams only against similar roles. The query in this guide uses the 2 hour variant. Worklytics also suggests, in a post on collaboration overload, a healthy baseline of two to three focus blocks of 60 minutes or more per day for knowledge roles, and says many teams show none. For more on what breaks focus up, see how distractions and interruptions affect focus time.

Timeline of one workday from 8 am to 7 pm with colored blocks for focused time, email, chat, and meetings, and shaded areas showing gaps between them.
Sample day view: meetings, email, and chat break up the time available for focused work. Source: Worklytics, 6 KPIs to Make Hybrid Work a Success.

Step 5: Score Connection and Meeting Load

More collaboration is not better, and collaboration hours alone say little. A more useful signal is whether people have close working relationships. This guide's connection score is the share of person-weeks in which a person had five or more strong collaborators, the isolation threshold from the table above. It needs no arbitrary hour targets. Also watch the other end: Worklytics reports that 12 or more strong collaborators tends to go with higher burnout risk, so a team can be connected and still carry too much overhead. To see when collaboration turns into overload, read how to prevent burnout by analyzing collaboration overload.

Meeting load is meeting hours in a week divided by five times the weekday span, so a person with 10 meeting hours and a 9 hour span has a load of about 22%. Dividing by span instead of a flat 40 hours keeps long-day and short-day workers comparable. Which meeting hours you use matters, as the definitions table shows, so check whether your metric counts direct time or exclusive meeting time. Because no threshold has been published for meeting load, the score uses the benchmarks approach: the best value is the 25th percentile and the worst is the 75th percentile of meeting load across everyone in the scored data. Calendar metadata can also show meeting quality signals: Viva Insights reports hours in large and long meetings, recurring meeting hours, meetings that end on time, conflicting meetings, and meetings scheduled with six hours or less of notice. To see how these work in practice, read how meeting insights reduce meeting overload.

Step 6: Convert the Metrics to Scores and Weight Them

Hours and percentages cannot be added together until they share a scale. This guide uses two methods. For continuous metrics, score linearly between a best and a worst value: for a metric where lower is better, such as meeting load, the score is (worst value minus your value) divided by (worst value minus best value), times 100. For a metric where higher is better, such as focus time, the score is (your value minus worst value) divided by (best value minus worst value), times 100. Keep every score between 0 and 100. For metrics with a bright line, such as a 9 hour day, score the share of person-weeks that stay on the right side of it.

ComponentWeightHow it is scoredWhere the cut-off comes from
Focus time30Linear, 0.5 hours a day scores 0 and 3.5 hours scores 100Worklytics research (3.5 hours). The floor is a choice
Meeting load25Linear, the 75th percentile scores 0 and the 25th percentile scores 100Your own data, the middle 50%
Connection20Share of person-weeks with 5 or more strong collaboratorsWorklytics isolation signal
Balance15Share of person-weeks with a weekday span of 9 hours or less and weekend work of 2 hours or lessWorklytics burnout indicators

The weights are a sample starting point, not an industry standard. They add up to 90, and the query divides by the total so the score still runs from 0 to 100. If you have an outcome measure from your own systems, such as on-time delivery, add it as a fifth component, because Microsoft 365 data does not measure output quality. Agree on the weights with managers and employee representatives, change them only at set times such as once a quarter, and record every change so that trends stay comparable. The next section shows how much the weights matter.

Calculate the Score in SQL

The query below computes team scores from Worklytics weekly aggregates using the real metric keys. It applies the active and 4 hour filters from Worklytics's sample SQL, flags each person-week against the two burnout thresholds, finds the meeting load range from your own data, drops groups under the minimum size, scores each component, applies the weights, and re-scores under shifted weights to test whether the ranking holds. It is written in BigQuery-style SQL, like Worklytics's own samples, and every threshold and weight sits in one params block.

-- Team productivity score from Worklytics weekly aggregates (BigQuery-style SQL)
-- Tables use the same placeholders as Worklytics' sample SQL:
--   {weekly_individual_aggregates}  columns: week, employeeId, key, value
--   {weekly_individual_groups}      columns: week, employeeId, active, custom_group_1, line_of_business

WITH params AS (            -- every threshold and weight lives here, so you can change them in one place
  SELECT
    3.5 AS focus_best,  0.5 AS focus_worst,   -- focus hours per day (3.5 h is the Worklytics research threshold)
    9.0 AS span_max,                          -- weekday span over 9 h is a Worklytics burnout indicator
    2.0 AS weekend_max,                       -- weekend work over 2 h is a Worklytics burnout indicator
    5 AS strong_min,                          -- fewer than 5 strong collaborators is Worklytics' isolation signal
    5 AS min_group_size,                      -- Microsoft's minimum for Viva Insights; raise it if you prefer
    30 AS w_focus, 25 AS w_meet, 20 AS w_collab, 15 AS w_balance
),
people AS (                 -- one row per active person per week
  SELECT
    a.week, a.employeeId, g.line_of_business AS team,
    MAX(CASE WHEN a.key = 'worklytics:weekdays:avg:timespan:hours' THEN a.value END) AS span_hours,
    MAX(CASE WHEN a.key = 'calendar:events:hours:meetings' THEN a.value END) AS meeting_hours,
    MAX(CASE WHEN a.key = 'worklytics:hours:in:focus:blocks:v3_5:flow' THEN a.value END) AS focus_hours,
    COALESCE(MAX(CASE WHEN a.key = 'worklytics:weekends:total:time:worked:hours' THEN a.value END), 0) AS weekend_hours,
    MAX(CASE WHEN a.key = 'collaborators:strong_count_distinct' THEN a.value END) AS strong_collabs
  FROM {weekly_individual_aggregates} a
  JOIN {weekly_individual_groups} g
    ON a.employeeId = g.employeeId AND a.week = g.week
  WHERE g.week >= '2026-07-06' AND g.week < '2026-09-01'   -- your scoring window
    AND g.active = true
    AND g.custom_group_1 IN ('Full-Time')                   -- exact filter varies by company
  GROUP BY a.week, a.employeeId, g.line_of_business
  HAVING MAX(CASE WHEN a.key = 'worklytics:weekdays:avg:timespan:hours' THEN a.value END) >= 4   -- same rule as Worklytics' sample SQL
),
person_weeks AS (           -- flag each person-week against the two burnout thresholds
  SELECT p.*,
    p.meeting_hours / (5.0 * p.span_hours) AS meeting_load,
    CASE WHEN p.span_hours > k.span_max OR p.weekend_hours > k.weekend_max THEN 0.0 ELSE 1.0 END AS within_limits
  FROM people p CROSS JOIN params k
  WHERE p.meeting_hours IS NOT NULL AND p.focus_hours IS NOT NULL
),
load_order AS (
  SELECT meeting_load, ROW_NUMBER() OVER (ORDER BY meeting_load) AS rn, COUNT(*) OVER () AS n
  FROM person_weeks
),
normal_range AS (           -- the middle 50% of meeting load across everyone scored: best = 25th percentile, worst = 75th
  SELECT MAX(CASE WHEN rn <= 0.25 * n THEN meeting_load END) AS meet_best,
         MAX(CASE WHEN rn <= 0.75 * n THEN meeting_load END) AS meet_worst
  FROM load_order
),
team_weeks AS (             -- one row per team per week; small groups are dropped here
  SELECT
    w.week, w.team, COUNT(DISTINCT w.employeeId) AS people,
    AVG(w.meeting_load) AS meeting_load,
    AVG(w.focus_hours) AS focus_hours,
    AVG(w.within_limits) AS share_within_limits,
    AVG(CASE WHEN w.strong_collabs >= (SELECT strong_min FROM params) THEN 1.0 ELSE 0.0 END) AS share_connected
  FROM person_weeks w
  GROUP BY w.week, w.team
  HAVING COUNT(DISTINCT w.employeeId) >= (SELECT min_group_size FROM params)
),
team_avg AS (               -- average each metric over the window
  SELECT team, COUNT(*) AS weeks_scored, ROUND(AVG(people), 0) AS avg_people,
         AVG(meeting_load) AS meeting_load, AVG(focus_hours) AS focus_hours,
         AVG(share_within_limits) AS share_within_limits, AVG(share_connected) AS share_connected
  FROM team_weeks
  GROUP BY team
),
components AS (             -- convert each metric to 0-100 (linear between worst and best, clipped)
  SELECT t.*, r.meet_best, r.meet_worst,
    CASE WHEN focus_hours >= p.focus_best THEN 100 WHEN focus_hours <= p.focus_worst THEN 0
         ELSE 100.0 * (focus_hours - p.focus_worst) / (p.focus_best - p.focus_worst) END AS focus_score,
    CASE WHEN meeting_load <= r.meet_best THEN 100 WHEN meeting_load >= r.meet_worst THEN 0
         ELSE 100.0 * (r.meet_worst - meeting_load) / (r.meet_worst - r.meet_best) END AS meeting_score,
    100.0 * share_connected AS collab_score,
    100.0 * share_within_limits AS balance_score
  FROM team_avg t CROSS JOIN params p CROSS JOIN normal_range r
),
scenarios AS (              -- base weights plus three shifts of 5 points, for the stability test
  SELECT 'base' AS scenario, w_focus, w_meet, w_collab, w_balance FROM params
  UNION ALL SELECT 'focus +5, meetings -5', w_focus + 5, w_meet - 5, w_collab, w_balance FROM params
  UNION ALL SELECT 'focus -5, meetings +5', w_focus - 5, w_meet + 5, w_collab, w_balance FROM params
  UNION ALL SELECT 'connection -5, meetings +5', w_focus, w_meet + 5, w_collab - 5, w_balance FROM params
),
scored AS (
  SELECT c.team, s.scenario,
    (s.w_focus * c.focus_score + s.w_meet * c.meeting_score + s.w_collab * c.collab_score + s.w_balance * c.balance_score)
      / (s.w_focus + s.w_meet + s.w_collab + s.w_balance) AS score
  FROM components c CROSS JOIN scenarios s
),
ranked AS (
  SELECT team, scenario, score, RANK() OVER (PARTITION BY scenario ORDER BY score DESC) AS rnk FROM scored
)
SELECT
  c.team, c.avg_people AS people,
  ROUND(c.focus_hours, 3) AS focus_hours_per_day, ROUND(100 * c.meeting_load, 2) AS meeting_pct_of_week,
  ROUND(100 * c.share_connected) AS pct_with_5plus_strong_collabs, ROUND(100 * c.share_within_limits) AS pct_within_limits,
  ROUND(c.focus_score, 1) AS focus, ROUND(c.meeting_score, 1) AS meetings, ROUND(c.collab_score, 1) AS connection, ROUND(c.balance_score, 1) AS balance,
  ROUND(100 * c.meet_best, 2) AS meeting_best_pct, ROUND(100 * c.meet_worst, 2) AS meeting_worst_pct,
  ROUND(MAX(CASE WHEN r.scenario = 'base' THEN r.score END), 1) AS score,
  MAX(CASE WHEN r.scenario = 'base' THEN r.rnk END) AS rank_base,
  MAX(CASE WHEN r.scenario = 'focus +5, meetings -5' THEN r.rnk END) AS rank_focus_up,
  MAX(CASE WHEN r.scenario = 'focus -5, meetings +5' THEN r.rnk END) AS rank_focus_down,
  MAX(CASE WHEN r.scenario = 'connection -5, meetings +5' THEN r.rnk END) AS rank_connection_down
FROM components c
JOIN ranked r ON r.team = c.team
GROUP BY c.team, c.avg_people, c.focus_hours, c.meeting_load, c.share_connected, c.share_within_limits,
         c.focus_score, c.meeting_score, c.collab_score, c.balance_score, c.meet_best, c.meet_worst
ORDER BY score DESC;

What the query returned on test data

The test dataset is synthetic. A script generated 57 people in six teams over eight weeks, with one part-timer per team and a few random inactive weeks. It is not Worklytics customer data, and the scores say nothing about real teams. Its purpose is to show that the query runs, that the filters work, and how to read the output. The four-person Legal team does not appear in the results because it falls under the minimum group size, and the part-timers are excluded.

TeamPeopleFocus hours a dayMeetings, % of weekWith 5+ strong collaborators, %Within limits, %Score
Engineering132.37619.31758579.5
Finance71.82725.87708353.7
Support80.99921.2604451.2
Marketing81.7726.01825049.7
Sales101.32932.07311117.9

Take the Support team. Its people averaged 0.999 focus hours a day, which scores (0.999 minus 0.5) divided by (3.5 minus 0.5), times 100, or 16.6. Its meetings take 21.2% of the workweek. In this data the meeting load range runs from 20.3% (best) to 28.74% (worst), so the score is (28.74 minus 21.2) divided by (28.74 minus 20.3), times 100, or about 89. Sixty percent of its person-weeks had five or more strong collaborators, a connection score of 60.4, and 44% stayed within the 9 hour and 2 hour limits, a balance score of 44.2. With weights of 30, 25, 20, and 15, the score is (30 times 16.6 plus 25 times 89.4 plus 20 times 60.4 plus 15 times 44.2) divided by 90, or 51.2. The query output rounds these numbers, so a hand calculation can differ by a tenth of a point.

Bar chart showing how four weighted components add up to a productivity score of 51.2 out of 100 for the Support team in the synthetic test data.
Support team from the sample query. Synthetic test data.

The weight-shift test is the useful part. It re-scores every team with 5 points of weight moved between focus and meetings, or between connection and meetings, and shows the rank each time.

TeamScoreRankRank if focus +5, meetings -5Rank if focus -5, meetings +5Rank if connection -5, meetings +5
Engineering79.51111
Finance53.72233
Support51.23422
Marketing49.74344
Sales17.95555

Three things stand out. First, Engineering stays first and Sales stays last under every weight setting, so those two placements are safe to report. Second, the three teams in the middle (Finance 53.7, Support 51.2, Marketing 49.7) sit within 4 points of each other and change places. Support drops to fourth when focus gains weight, because it has the weakest focus score of the three, and rises to second when focus loses weight, because it has the strongest meeting score of the three. Third, the meeting range is narrow. The middle 50% of person-weeks spans only about 8 points of the workweek, so Support at 21.2% scores 89 while Marketing at 26.0% scores 32. A gap of under 5 points in meeting time became a gap of 57 points in the score. If your own range is this tight, widen it (for example, use the 10th and 90th percentiles) or score meetings against a target, and rerun the weight test.

The practical rule is to report teams in bands, such as top, middle, and bottom, and to save exact numbers for trends within one team. Here that would put Engineering on top, Sales at the bottom, and the other three together in the middle.

Step 7: Show the Score With Context

A bare number invites arguments. Worklytics's Flexible Work Scorecard shows one way to present a score: each measure gets a letter grade and a rank against peers. Build separate views for separate readers, starting from Microsoft's Ways of working Power BI report if you use Viva Insights. Leaders see score trends and department comparisons in bands. Managers see their own team's components. Individuals see personal insights, private to them, and never a ranking or a peer comparison. For tool options, see best employee productivity dashboards, or see how the workplace insights dashboard presents these metrics from several Microsoft 365 sources.

Sample Flexible Work Scorecard with an overall B+ grade and separate letter grades and peer rankings for engagement and attendance measures.
Sample Flexible Work Scorecard: each measure gets a grade and a rank against peers. Source: Worklytics.

A 12-Week Rollout Plan

Twelve weeks is a planning estimate, not a benchmark.

WeeksWhat you doWhat you have at the end
1 to 2Get admin approval, choose a data route, set the minimum group size, and tell employees what is measured.A signed-off data plan
3 to 4Clean the data, apply the active and exclusion filters, and record a baseline for every metric.Clean data and baseline numbers
5 to 8Choose the focus variant for each role, build the focus, connection, and meeting load scores, and check the meeting range.Tested component scores
9 to 10Set weights, run the weight test, and decide on bands.A stable scoring rule
11 to 12Build the manager and leader views, set access rules, and review with employee representatives.A pilot dashboard and written rules

Using Google Workspace Instead of Microsoft 365?

The scoring method works the same way with Google Workspace data. Only the data source changes: Google Calendar, Gmail metadata, Drive activity, and Google Meet take the place of Outlook, Exchange, SharePoint, and Teams. Google lets administrators export Workspace logs for Calendar, Gmail, Drive, Meet, and other services to BigQuery, where you can calculate the same metrics. The export is limited to certain editions, such as Enterprise Standard and Plus and Education Standard and Plus, so check your edition first. For setup details, see the Worklytics pages on Google Workspace analytics, Google Calendar analytics, and Google Meet analytics.

What Usually Goes Wrong

Most scoring projects fail on definitions and interpretation, not on math. These are the problems to check first:

  • Overlapping meetings are double counted. Direct meeting time adds every meeting's hours, so an overbooked calendar looks busier than it is. Use meeting count or exclusive meeting time.
  • One focus block length is used for every role. Sales and customer success teams score poorly on a 2 hour block even when their days are healthy.
  • Teams a few points apart get ranked. In the test data, three teams within 4 points changed places when 5 points of weight moved. Report bands, not ranks.
  • The meeting range is too narrow. When the middle 50% covers only a few points of the workweek, small differences swing the score from 0 to 100. Check the spread before you trust the meeting score.
  • Published thresholds are treated as facts about your company. The 9 hour, 2 hour, 3.5 hour, and 5 collaborator lines come from Worklytics posts with no published sample size. Test them against your own survey data.
  • Small teams disappear. A minimum group size drops them from the results. Decide on a roll-up rule before you publish.
  • Weights change without a record. Trends become impossible to read. Log every change and change weights only at set times.
  • The score is used as a performance rating. Scores work at team level. Individual rankings hurt trust and go against the privacy rules above, and a score based on activity cannot tell you whether the work was good. Pair it with an outcome measure.
  • Employees are left out of the conversation. Publish what is measured, show each team its own data before leaders see comparisons, and ask employee representatives to review the metrics and weights. Check for differences by role, location, or working pattern, because part-time staff or people in other time zones can look worse for reasons unrelated to work.
  • Data arrives later than expected. Viva Insights refreshes weekly, and if you export to Microsoft Fabric you should schedule the refresh after the weekly Viva Insights update. Do not expect daily numbers.

There is no published standard for how fast results appear. Judge the program by whether managers use the numbers and whether focus time and meeting load move against each team's own baseline over several months.

Conclusion

A productivity score from Microsoft 365 data is only as good as its definitions and cut-offs. Decide what counts as active, as a meeting, and as focus time, use published thresholds where they exist and your own data where they do not, and test whether the ranking survives a change in weights. Balance analytical rigor with privacy protection, so that measurement supports employees rather than simply monitoring them. Start small: pick a few metrics, set a baseline, report by team, and adjust as you learn.

Worklytics provides a privacy-first alternative to manual models, connecting to Microsoft 365 data while keeping message content out of scope. See how the Microsoft 365 integration works, or use this guide to build the score yourself first.

Frequently Asked Questions

Is Microsoft Productivity Score the same as Adoption Score?

Microsoft's Productivity Score is now called Adoption Score in the Microsoft 365 admin center. It measures how an organization uses Microsoft 365, not how productive individual employees are.

Can managers see an individual employee's Microsoft productivity score?

No. Microsoft says Adoption Score insights are calculated at the organization level, and no one in an organization can use it to see how an individual uses Microsoft 365. Viva Insights shows managers team-level insights and requires a minimum group size of at least five.

What is a good productivity score?

There is no universal number for a custom score, because weights and cut-offs are your own choices. Compare each team with its own baseline and with similar teams, and avoid ranking teams that are only a few points apart. For Adoption Score, Microsoft says many mature organizations operate effectively at 80 to 90 points per category, so a perfect score is not the goal.

Which thresholds should I use for focus time and workday length?

Worklytics has published 3.5 hours of daily focus time, a workday of no more than 9 hours, weekend work of no more than 2 hours, and at least 5 strong collaborators as useful lines. Treat them as starting points, because the posts do not publish sample sizes. Check them against your own survey data and write down which ones you used.

How do I calculate the score from Worklytics data?

Use the weekly aggregates export, which has one row per person, week, and metric. Filter to active people, flag each person-week against the workday and weekend thresholds, score focus time, meeting load, connection, and balance, drop groups under your minimum size, and apply weights. The SQL query above does this with the real metric keys and tests whether the ranking holds when the weights shift.

Why do two tools report different meeting hours for the same team?

They define the numbers differently. Tools differ on what counts as a meeting, whether declined invites count, and whether overlapping meetings are counted once. Worklytics, for example, counts events with at least two human attendees who have not declined, and offers both direct time and exclusive meeting time. Pick one definition before you score.

Request a demo

Schedule a demo with our team to learn how Worklytics can help your organization.

Book a Demo