Reports
Seven reports over one filter, and a way to keep a filter you use often.
The shape of it
Every report accepts the same filter: period, projects, types, statuses, priorities, specific people and agents, and whether to count humans, agents or both. Reports differ in what they aggregate, never in what they can be narrowed by.
That is deliberate. Seven reports with seven filter implementations is seven places for "last quarter" to come to mean something slightly different.
Weeks ahead is the exception, and says so on the page. It looks forward, so a period is meaningless to it: it takes who, and how far ahead, and nothing else. A project or a status is a fact about issues, and that report is about calendars.
Two entries on the index are not filtered reports at all but pages of their own — the Timeline and the Gantt — and they are listed there because that is where somebody looks for them.
What each one answers
| Report | The question |
|---|---|
| Status report | Where are we, what moved, what is stuck, what is due |
| Time report | Where the hours went, and whose they were |
| Throughput | Are we keeping up, and how long does work take |
| Flow | Where the time actually went — how much was work, how much was waiting, and how much the board only claimed was work |
| Workload | Who is carrying what, right now |
| Weeks ahead | What each week ahead actually holds, with leave and public holidays taken out |
| Agent performance | Which agents are working, and which are failing |
| Timeline | Which epics land when, and which will not — see Timeline |
| Gantt | The plan against what happened — not for Management |
| Standup | Who is behind on what — owners, admins and management |
Status report
Built for somebody who is not in the tool. Overdue and on-hold come first, because a status report is read for its problems.
Closings are read from the activity log rather than from updated_at: an issue
somebody comments on next week would otherwise slide out of the period it was
actually finished in.
Time report
Human time and machine time in the same table, split but never separated. An agent's twenty minutes is twenty minutes the team did not spend, and a report that hides it understates both the cost and the help.
Estimates are only compared where an estimate exists. An absent estimate is not a promise, and counting it as zero would invent an overrun nobody agreed to.
How much of it was measured. A headline figure says what proportion of the total was watched happening rather than remembered afterwards:
Measured 13%
1h 35m observed · 1d 2h 32m recalled
A stopwatch knows when it started and a daemon knows when a run began and ended. Somebody typing "about three hours" into a form on a Friday afternoon does not, and a total that adds the three together without saying so is a guess wearing a number's clothes.
Measured how filters the report to one provenance. It is offered here and
nowhere else: a source belongs to an hour, not to an issue, and a control that is
displayed and then ignored is worse than one that is absent. ?source= on
GET /time-entries does the same over the API.
How round the recalled figures are. Memory lands on the half hour: somebody typing in Friday's hours writes 2h, not 1h 50m, and a stopwatch never does because it does not know what a round number is. So the report says what share of recalled entries land on a multiple of thirty minutes beside the share of observed ones that do — 80% of recalled entries land on a half-hour; 10% of observed ones do. It is a measurement of recall bias the timesheet has always contained. About the figures, never about a person: a whole team recalls in round numbers, and the point is knowing how much of a report rests on them. Withheld under ten entries of either kind, because four recalled entries, three of them round, is 75% and means nothing.
Throughput
Opened against closed, week by week. Whether the backlog is growing is a fact no board will ever show you, because a board only knows the present.
Cycle time is a median. One issue that sat in the backlog for a year should not define the team's reputation. It is measured from the first status somebody works in to the day the work closed, and it is Flow's figure — read from there rather than computed again, so the decomposition on one page adds up to the number on the other.
Did it land when we said?
The trust number, and the one nothing in the product had ever measured. Every projection here — the Gantt's bars and its capacity rule, the routing flyout's dates — is an argument about the future, and a schedule nobody has audited is a schedule people quietly stop reading.
Four figures: the share that landed on the day or before, the share within two days either side, the median slip and the 85th percentile.
Measured against the due date, which is a deliberate limit rather than a shortcut. A projection is a function of the capacity on the day it was made, and none of that is kept — so recomputing one for a finished issue would produce a date today's calendar implies and yesterday's never said. The due date is the only forecast that was written down. That makes this the other half of Promise this in the routing flyout: committing a projection is what turns this figure from did somebody guess a Friday correctly into does our schedule hold.
The slip is signed, so early and late do not cancel. A team that is systematically early has a different problem from one that is systematically late, and an absolute figure reports them as the same team.
Work finished with no date is counted apart, never as a success. It is usually the larger population, and folding it in either direction would be the whole finding.
Unfinished work is left out. An issue past its date and still open is late, which the board and the standup already say — it is not a measurement of whether a promise held.
Flow
Every tracker reports how long work takes. Almost none say how much of that was work, because almost none keep the status history to answer it — and the answer is usually that most of the time was nobody doing anything.
Lead time 15.5d · working 41%, stalled 6%, waiting 53% · of the waiting, 24 days in Backlog and 7 in Todo
"Eleven days, of which three were worked" is a different conversation from "eleven days", and it is the only one that points anywhere: you cannot make the three much faster, and the eight are often a decision.
Two figures, each with its tail. Lead time is opened to closed — the wait somebody outside the team experienced. Cycle time starts when the work did. Beside each is what 85% finish within, because a median is no basis for a promise: "half finish within four days" is not something to tell anybody, where "85% within twelve" is.
Working and waiting are your own classification, not a list of names in the
report. A status counts as working when the workspace classifies it as
started and does not mark it paused — exactly what in flight already
means in the status report. So the lever is one you already have: a team that
thinks a review queue is not work marks that status paused, and its time moves.
Nothing in the report changes.
Stalled is the third figure, and the honest half of working. A started status is counted as work — and an issue that sat in one untouched for a week was not being worked, it was a question about who was doing it. So the time in working statuses is split again: the days nobody touched the issue beyond a three-day allowance are stalled, and working is what is left.
Working 47% · Stalled 12% · Waiting 41%
"Touched" means any activity of any kind, or a day with hours logged — not a status change, which is what the segments are cut on, but evidence the person who had it was there. The first three days of any silence are the ordinary rhythm of work and are not held against anybody; only what runs past them counts. Waiting cannot stall: nobody was meant to be on it.
Stalled comes out of working, not out of waiting, so the three read apart: time in a queue, time being worked, and time the board claimed was being worked.
The per-status table is the actionable half. "Waiting 8 days" says something is wrong; "6 of the 8 in review" says what.
People and machines, side by side. Machines do not wait: an agent's work sits in a queue for minutes and a person's for days. Split by what is doing the work rather than by who — this is an argument about what to route where, not a league table.
Calendar days, weekends included, and this is the only report here that does not count working days. Capacity is about how much somebody can do, so a weekend is not capacity. This is about how long a thing sat, and a thing that sat over a weekend sat for two more days. Nobody waiting on it experienced a four-day week.
Work with no recorded history is counted apart, never as a zero. An issue closed before any of this existed has a lead time and no breakdown. Filing its whole life under whatever status it is in now would invent a figure; dropping it would make the lead time a median of the recent half. So it is in the lead time, out of the split, and the page says how many. Below half coverage the split is not drawn at all — a breakdown from a fifth of the work, presented as the shape of the whole, is worse than no breakdown.
Only finished work. An unfinished issue has no lead time yet, only a wait that is still running.
Workload
Deliberately not filtered by period — what somebody had on their plate last March is not a useful question.
Issues without an estimate are counted and shown rather than averaged away. A plate with three unestimated issues on it is not a light plate.
Weeks ahead
The only report here that looks forward, and the same capacity model as everything else — pointed the other way. Absence reduces what a week holds, so ask what the weeks ahead hold before promising one of them.
The case it exists for is the one nobody catches in time. Easter and the week between Christmas and New Year are systematically thin — two or three working days, half the team away — and a cycle planned across one of them is short before it starts.
Three things make it an answer rather than a calendar with numbers on it:
A public holiday counts whether or not anybody has logged it. A report that read only the timesheet would show Holy Week as a full week until somebody entered the leave, which is precisely the moment it stops being useful. So the holidays are computed as well as read — and counted once, never twice, because the computed subtraction only ever takes what the recorded one left.
Machines do not take holidays. People and agents share one capacity model here, so a thin week is thin for one half of the workforce and not the other. Week 52 reading 34% for people and a full week for the machines is a routing argument, and it is invisible in any tool that models only people.
The week in progress is never called thin. Two days left of five on a Thursday is a fact about it being Thursday. A warning that appears every week from Wednesday onwards is a warning nobody reads.
Everything is measured against each actor's own week. Somebody on three days is not a permanently thin week; they are a full week of three days.
The one thing it cannot see is leave nobody has entered yet. It says so on the page.
Agent performance
Success rate counts finished runs only. A run still in flight is not yet a success or a failure, and counting it as either is a lie that flatters.
Standup
Built for the Monday morning meeting: every open issue on one timeline, grouped by whoever is carrying it, with anything past its estimate marked.
Sorted worst first — most past estimate, then most overdue, then most work. A list in alphabetical order makes everybody read all of it to find the two rows the meeting is about.
Each person's header is meant to be readable on its own: eight issues, two past estimate, 14h of 9h. The hours count only the estimated issues, for the same reason the cost report separates unpriced hours from free ones — an unestimated issue that has taken ten hours would otherwise make somebody look under budget by ten hours.
Every marked row carries both numbers, not just the colour. "Past estimate" is an argument; five hours against two is its evidence.
Owners, admins and management. This is one of two restricted entries — the Gantt is the other, because it is a page that edits dates rather than a report that reads them — and the restriction is the point: it says who is behind. Being able to read that is the same kind of trust as deciding who works here, and deliberately not the same as running the agent workforce — a developer gets the second without the first, which is the split the roles already draw. The tile is absent for everybody else rather than present and refusing, and the CSV export is gated with it.
It groups by assignee, not by the developer role. Those are different things and the role would have been the wrong one: in the workspace this was first asked for, everybody holding an issue was an owner or an admin, and the four people with the developer role had none. Grouping by role would have shown four empty rows and hidden all the work.
Somebody with nothing open simply has no row. That is a fact the meeting wants — it is either good news or somebody who has been forgotten — but inventing a row to say "nothing" is not this report's job.
Untouched. Beside past estimate and overdue, a third badge: how many of a person's started issues nobody has touched for three days or more, and on each such row how long — untouched 6d. "In progress" used to hide this; a started card read as work happening whether or not anything was. The same badge appears on the board card itself. See Concepts — started and untouched.
Time has to be attributable
issues.spent_minutes is a cached total. Underneath it sits one entry per piece
of work, each with an actor, a date and a source:
- A person logs time from the issue's properties panel — source
manual - The stopwatch writes one when it is stopped — source
timer - An agent writes an entry when its run completes — source
agent_run
timer and agent_run are measured: something was running while the work
happened. manual is recalled. Hours the timer recorded before v0.26.0 still
read as manual — they were not reclassified, because guessing which of them
came from a stopwatch would be inventing data.
The date used everywhere is when the work happened, not when it was entered. An hour worked on Monday and logged on Friday belongs to Monday.
Agent entries cannot be deleted. They are a record of something that happened, not a claim somebody typed, and removing one would make the run history disagree with the time report.
Saved views
Narrow a report, then Save this view. It appears on the reports index and in the picker on that report.
Saved views are private by default. A saved filter is a working note until its author decides it is a team artefact, so sharing is a deliberate act rather than the fallback. A shared view can be opened by anyone in the workspace and changed only by whoever made it.
A view that has been deleted or unshared stops applying rather than breaking the page somebody bookmarked.
The issue board saves views the same way and out of the same drawer — narrow it, then Filter → Save this view. Same rules, same table: private until shared, yours to delete, and named in the URL so a link opens what you were looking at.
Each view is listed where it was saved. A board filter appears on the board and a report's filter on the reports index — the record has always said which surface it came from, and every list asks. A board filter was briefly listed here too, under a link that led nowhere, because no report answers to the board.
The list also respects who may open what: a saved Standup view is not listed for somebody who cannot open the standup.
How they are computed
Aggregation happens in the database. The first version fetched every row in the period and grouped in PHP, which is fine for a fortnight and is a memory limit for All time — the one option in the picker nobody thinks twice about choosing.
Three things deliberately stay in PHP:
- Medians, which have no portable SQL. Only the two timestamp columns are read, never whole models.
- Weekly buckets, because grouping by week means
YEARWEEKon MySQL andstrftimeon SQLite. Days are grouped in SQL and bucketed in PHP: a year is at most 365 rows. - Name lookups, resolved in two queries rather than one per row.
The lists on a status report are capped at 200 rows, since the page shows twenty-five. The export builds its own and is not capped.
Every aggregate is exercised against both MySQL and SQLite in the test suite,
because conditional sums, DATE() and JSON paths are exactly where the two engines
diverge — an accidental dependence on one of them is cheapest to find by running
against the other. It is not a claim that either will do: Félagi requires MySQL 8,
and SQLite is a test target.
Export
Every report exports as CSV, honouring the filters on screen — and a report that is restricted on screen is restricted here too. The page and the export are two doors to the same data, and the export is a plain link with the filters in the query string: the easier of the two to reach by accident.
CSV rather than PDF because a spreadsheet is what people actually do with a report: re-sort it, total a column, paste it into something else. A PDF looks finished and can only be read.
The time export is one row per entry rather than a summary, so whoever opens it can group it whichever way they came for. It is streamed lazily — each row is written and forgotten, so a year of entries is never resident at once.
Exports lead with a byte-order mark. Without one, Excel reads UTF-8 as Latin-1 and turns Félagi into Félagi on the first row of every file.
How estimates behave
At the bottom of the time report: the factor between what this workspace estimates and what the work takes.
An estimate is usually not written by the person who does the work — a lead sizes a backlog, a planning meeting agrees a figure, somebody types "2h" because the last one like it was 2h. So this measures the estimating, not the estimator, and the place to spend it is a plan rather than a person.
What counts as evidence
Three conditions, each of which throws data away on purpose:
| It has an estimate | Obviously |
| Every hour on it was measured | A stopwatch or a daemon. One remembered entry and the actual is partly a recollection, which makes the ratio partly fiction |
| All that time is one actor's | An estimate is for the work, so an issue two people shared says nothing about either of their estimating |
That last pair is the whole design. Félagi already separates measured hours from remembered ones, and calibrating an estimate against remembered time compares a guess with a guess and calls the ratio a fact.
It also has a consequence worth knowing: agents end up the best-calibrated actors in the product, and not because they are better estimators. Every minute a daemon reports was watched, so their column has no recollection in it at all. It is the one number here no other tracker can compute, because no other tracker employs the machine.
The number
The median of the per-issue ratios, not the ratio of the totals. Totals let one large issue decide the factor — a fortnight-long task that ran over would drown twenty small ones that were fine. The question is how wrong a typical estimate is, so every estimate gets one vote.
Below five measured issues it says nothing. A ratio from two observations is a coincidence, and presenting it as a factor is worse than silence.
A factor is capped at 4×. That is not a correction of the data but a refusal to multiply by it: a factor of twelve says estimates and reality are unrelated here, and the resulting number is one nobody would use.
Somebody with too little history of their own gets the workspace's factor, not 1.0. A new colleague has no history and is not therefore a perfect estimator.
It is not narrowed by the report's filters. Calibration is a property of a workspace's whole measured history, and a number that moves when you change a date picker is not a calibration.
Where it is spent
On a cycle's plan, beside its capacity:
7d 3h 30m planned has historically meant 8d 6h 13m
The case it catches that a capacity check cannot: a window that fits on paper and has never fitted in practice. When that happens the page says so.
Not there yet
- No scheduled reports. Nothing arrives in an inbox on a Monday morning.
- Cost is only as complete as your CLIs. Claude Code reports it; Elyra does not
yet, and a run that reported nothing shows
—rather than$0.00— a total says how many runs it covers, so a figure from three of forty cannot be multiplied out by mistake. - No currency other than USD, and no conversion. An agent hour is time rather than money.
- No per-project time budget, and no billable flag.
- Stalled is measured on finished work in Flow, and on open work everywhere else. The two use the same three-day allowance but are not the same figure: Flow says how much of a lead time was stalled; the board and standup say what is stalled now.
- No trend on Flow. It answers "where does the time go" for a period, not "is the waiting getting worse". Two periods side by side is the comparison, and it is one you make by changing the period twice.