In the autumn of 2023, Fabrizio Dell’Acqua and nine co-authors from Harvard, MIT, Wharton, Warwick and Boston Consulting Group published one of the most cleanly constructed field studies to date on the productivity effect of generative language models. 758 BCG consultants, a pre-registered between-subjects design, three groups, realistic consulting tasks — and, at the end, two figures that have since been quoted in almost every talk on AI in knowledge work: a speed gain of 25.1 per cent and a quality gain of more than 40 per cent for tasks inside what the authors call the "jagged frontier"[1][2]. For funding management the study matters for three reasons: it measures on a task portfolio structurally close to that of an application; it names the limits of the lever; and it supplies a concept — the jagged frontier — that explains why the same human-in-the-loop approach can produce brilliant and catastrophic results side by side on a funding application. This piece places the study's design, results and limits in context and carries them over to application work.
Study design and setup
Harvard Business School working paper 24-013 is titled "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality" and was published in September 2023[1]. The SSRN preprint version, paper ID 4573321, has been available since the same period[2]. In 2025 the peer-reviewed final version appeared in Organization Science[3] — it confirmed the preprint's core findings.
The experiment was pre-registered, meaning the sample and the hypotheses were fixed publicly before the data were collected. The participants were 758 consultants at Boston Consulting Group, recruited across seniority levels and offices. All first completed a baseline task of comparable difficulty, so that individual performance differences could be factored out. They were then assigned at random to three groups: a control group with no access to generative AI; a GPT-4 group with access to the model; and a GPT-4 group that first received a short introduction to prompt engineering[1][2].
The measurement itself took place on 18 realistic consulting tasks that the author team developed together with BCG. The task portfolio covers a consultant's day-to-day work broadly — creativity, analytical thinking, writing, persuading, client communication. Alongside the 18 tasks inside the presumed AI frontier, the team added a 19th that was deliberately constructed outside it: a quantitative analysis in which GPT-4 steered towards plausible but wrong conclusions[1]. This two-part cut — 18 inside, 1 outside — is methodologically decisive and comes up again in the section on the jagged frontier.
The quality of the submitted solutions was assessed by independent graders, blind to the treatment assignment. Speed was captured through time spent per task, quantity through the number of tasks completed per time window. The design therefore allows three independent effect estimates: on speed, on quality and on throughput[2].
The two levers: speed and quality
The study's most widely quoted figures come from the measurement inside the frontier. Consultants with access to GPT-4 completed tasks on average 25.1 per cent faster, produced results rated more than 40 per cent higher on human quality assessment, and completed 12.2 per cent more tasks per time window[1][2]. The effect sizes were statistically significant for all three levers and measurable in both AI groups — with and without the prompt engineering overview.
The finding on the distribution of performance is especially striking. The study shows that the productivity gains were not evenly spread: consultants from the bottom third of baseline performance benefited relatively more than those from the top third. AI therefore acts as a levelling lever, not as an amplifier of individual excellence. This pattern matches the findings of two other large field studies from the same period: in 2023, Noy and Zhang measured, in Science, on a sample of 444 college-educated knowledge workers on writing tasks, a time saving of 0.8 standard deviations and a quality increase of 0.4 standard deviations, with weaker writers benefiting disproportionately[4]. Brynjolfsson, Li and Raymond report an average productivity gain of 14 per cent for 5,179 call centre agents, made up of 34 per cent among newcomers and practically no effect among experienced staff[5].
The parallel with Peng et al. on GitHub Copilot — where a controlled experiment with professional developers found a 55.8 per cent speed improvement on a defined programming task — rounds out the picture[6]. The magnitudes vary by task type, industry and model generation; the direction is robust across the studies. What matters for the argument that follows, though, is not whether a lever exists — but how far it reaches and where it ends.
The jagged frontier concept
The paper's central conceptual contribution is neither the figure 25.1 nor the figure 40. It is the concept of the jagged frontier — the ragged, irregular boundary between tasks where a large language model helps reliably and at high quality, and tasks where it produces plausible-sounding but wrong results[1]. The boundary does not run parallel to any intuitive axis of difficulty. Two tasks can look equally hard to a human worker — and yet one lies inside the AI's capability and the other outside. Dell’Acqua et al. show this empirically.
For the 18 tasks inside the frontier, the AI group produced the gains quoted above on average. For the one task outside it — a quantitative analysis with misleading figures in the brief — the AI group was 19 percentage points less likely to give the right answer than the control group. Here the AI did not merely fail to help its users; it pulled them systematically in the wrong direction[1][2]. The pattern refutes the assumption that human oversight will correct things when in doubt: the workers largely adopted the AI's output, because at first glance it looked coherent.
The practical conclusion is sobering. Anyone designing human-in-the-loop procedures cannot rely on a blanket productivity formula. They have to work out, for the specific field of work, which sub-tasks lie inside and which outside the frontier, and differentiate the division of labour between human and model accordingly. The OECD paper on the effects of generative AI on productivity, from June 2025, sums the point up briefly: the effect depends on task type and experience, and the quality of human-AI collaboration is the decisive parameter[8].
Limits of the study
Even the cleanest field study has limits, and Dell’Acqua et al. name several themselves. First, the sample is homogeneous: 758 consultants from the same firm, with comparable training and a uniform working culture. The results cannot be transferred automatically to knowledge work in funding application departments, research coordination or public administration.
Second, a single model generation was tested — GPT-4 in its summer 2023 version. Both the capability of the models and the nature of their errors have moved on since; the precise shape of the frontier is model-dependent and shifts with each generation.
Third, the tasks are short, self-contained and without extended context. A funding application is the opposite: an artefact spanning months, several dozen interwoven components, several layers of primary sources, formal deadlines. Studies on such long, institutionally embedded tasks are so far missing — the OECD points out explicitly that the bulk of the experimental evidence comes from short, isolated tasks[8].
Fourth, the study measures effects at the individual level. In his macroeconomic assessment, Acemoglu argues that microeconomic productivity gains of 25 or 40 per cent per task by no means imply an equally large effect for the economy as a whole. His task-based framework puts generative AI's contribution to total factor productivity over ten years at 0.66 per cent at most — because only some activities are exposed and the savings per activity are heavily diluted in aggregation[7]. That does not refute the Dell’Acqua study's finding, but it does sharpen its reach: 25 and 40 per cent on a single task are real, but they do not translate one-to-one into organisational productivity.
Fifth, and most relevant for the funding application: the study measures no reuse effects. Every task was worked on once; components were not reused across applications; institutions built up no store of experience. The study therefore leaves a second lever — the component used more than once — structurally unmeasured.
Transferring this to the funding application
A consultant's task list and an applicant's overlap in several places. Both write structured documents with formal requirements on structure, tone and strength of argument. Both draw on heterogeneous primary sources — market data, studies, legal texts, technical specifications — and have to build a consistent narrative from them. Both work under time pressure against an external addressee with its own assessment criteria. The gains measured on the Dell’Acqua 18-task battery should in principle be reproducible at many points in an application[1].
At the same time, a funding application differs from a BCG task in at least three dimensions. First, the logic of assessment: programmes such as ZIM or the research allowance operate with criteria that follow the OECD's Frascati Manual — novelty, uncertainty, systematic approach, transferability, creativity[9]. These criteria are not literary but tied to the facts of the case. A language model that formulates "sufficient uncertainty" in the application text says nothing about whether the uncertainty actually exists in the project. This is precisely a typical outside-frontier zone: the text sounds coherent, the underlying facts are not.
Second, the legal protection of the data being processed. Documents relevant to an application regularly contain client confidences, personnel data, technical internals. Section 203 of the Criminal Code penalises the disclosure of such secrets by members of certain professions and, in the context of professional confidentiality obligations, sets hard limits at the interface between the person doing the work and the model[10]. Which data a language model may see is therefore not a pure productivity question but a compliance question. In the funding context, the detailed path — which documents may be processed with which model on which infrastructure — is dealt with in a separate piece and linked there to the primary sources (see "ChatGPT and section 203 of the Criminal Code in the funding application").
Third, the lifespan of the artefact. An application is not submitted once; it runs through preliminary talks, interim versions, requests for further information, the grant notice, interim proofs of use and a final proof of use. At each of these steps, text components from the application are needed again — slightly varied, in a different context, for a different addressee. That is the reuse lever, the one the Dell’Acqua study does not measure and the one that, implemented cleanly on the platform side, produces greater effects in practice than the one-off speed gain on an isolated task.
The second order of magnitude on which funding applications differ from BCG tasks is the 20/80 asymmetry between generic and client-specific content. In a typical application, a considerable share of the text — the programme description, the legal framework, the form fields, the prescribed structure — is identical or nearly identical across applicants. The specific share — the project, the staff, the figures — is small but decisive for assessment. The human-in-the-loop lever acts above all on the generic share; the specific share stays human-checked. We treat this asymmetry in a separate piece as a design principle for application architectures.
What the HITL lever does not replace
The Dell’Acqua study shows what a well-instrumented human-in-the-loop approach can do, and just as precisely what it cannot. It replaces neither the technical assessment of the facts nor their placement within the legal and institutional framework of the funding programme in question. At every point where the task crosses the jagged frontier, the lever flips — from accelerator to risk amplifier[1].
For funding management, a set of practical consequences follows. First: the human and model roles have to be differentiated along the frontier. Standard components, structuring text, smoothing, deriving from existing primary sources — inside the frontier, this is where the lever works. Establishing the facts, assessing unevidenced primary sources, checking risk, coordinating with the funding body and the tax office — outside the frontier, this is where the human remains responsible. Second: the central data basis must consist of checked primary sources, not freely generated text. The OECD points out explicitly that quality gains from generative AI arise above all where the human side brings good source material and good oversight[8]. Third: every working trace has to remain auditable — which version, which component, which source, which person.
upsmart writes applications — and the work on them comes together on a platform where projects, documents and approvals hang together: versioned components, linked primary sources, checked relationships between projects, applications and proofs of use. The levers from the Dell’Acqua study are effective in the platform at the points where the frontier permits it — and the platform makes visible at the same time where the frontier is being crossed and human review remains necessary. The 25.1 per cent on speed and 40 per cent on quality are no free pass; they are an empirically evidenced upper bound for the inside-frontier field, realised only when the boundary is drawn precisely.
The real task, then, is not to feed more AI into application work, but to know the boundary between inside and outside for each individual sub-task and to cut the working role accordingly. The Dell’Acqua study supplies the vocabulary for that. Funding practice supplies the detailed topology. Between the two lies the space where a platform like upsmart makes the difference.
- [1]Dell'Acqua, F., McFowland III, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., Lakhani, K. R. — Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality. Harvard Business School Working Paper 24-013, September 2023.Harvard Business School, Faculty & Research · 2023Open source
- [2]Dell'Acqua et al. — Navigating the Jagged Technological Frontier (HBS Working Paper 24-013, direct PDF version).Harvard Business School, Publication Files · 2023Open source
- [3]Dell'Acqua et al. — Navigating the Jagged Technological Frontier. Preprint version, SSRN paper ID 4573321.Social Science Research Network (SSRN) · 2023Open source
- [4]Dell'Acqua et al. — Navigating the Jagged Technological Frontier. Organization Science, peer-reviewed published version.INFORMS — Organization Science · 2025Open source
- [5]Noy, S. & Zhang, W. — Experimental evidence on the productivity effects of generative artificial intelligence. Science, vol. 381, pp. 187–192, 14 July 2023. DOI 10.1126/science.adh2586.American Association for the Advancement of Science (AAAS), Science · 2023Open source
- [6]Brynjolfsson, E., Li, D., Raymond, L. — Generative AI at Work. NBER Working Paper 31161, April 2023 (extended 2024/2025 in The Quarterly Journal of Economics).National Bureau of Economic Research (NBER) · 2023Open source
- [7]Peng, S., Kalliamvakou, E., Cihon, P., Demirer, M. — The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590, February 2023.arXiv / Microsoft Research · 2023Open source
- [8]Acemoglu, D. — The Simple Macroeconomics of AI. NBER Working Paper 32487, May 2024 (published 2025 in Economic Policy, vol. 40).National Bureau of Economic Research (NBER) · 2024Open source
- [9]OECD — The effects of generative AI on productivity, innovation and entrepreneurship. OECD Science, Technology and Industry Policy Paper, June 2025. DOI 10.1787/b21df222-en.Organisation for Economic Co-operation and Development (OECD) · 2025Open source
- [10]OECD — Frascati Manual 2015: Guidelines for Collecting and Reporting Data on Research and Experimental Development. Chapter 2: Concepts and definitions for identifying R&D.OECD Publishing, Paris · 2015Open source
- [11]Section 203 of the German Criminal Code (StGB) — violation of private secrets.Federal Ministry of Justice, gesetze-im-internet.de · 2025Open source
The 20/80 asymmetry: why writing the application is the smallest part of the work
The FDP Faculty Burden Survey, the EC's Horizon Europe interim evaluation and the ECA error statistics show where the effort actually lies — and where classic consultants structurally cannot reach.
The reuse lever: how an application platform pays for itself from cycle 2–3
NIH R01 resubmission data show the empirical advantage of the second submission. What that means for the roadmap of German funding applications.
ChatGPT in the funding application: why it breaches section 203 of the Criminal Code
Tax advisers, lawyers and auditors risk criminal consequences when they use ChatGPT. What the DSK put on record about it in 2024.
From the analysis into the application.
We show the platform on a real case.