Abstract
This paper argues that a single structural mechanism, the substitution of measurement for exit, recurs across two apparently unrelated interventions separated by a decade and by a shift in technology: the Plan 2020 civic planning process in Indianapolis, conducted from 2014 to 2016, and Project Evident's 2026 research brief titled Scaling Impact with AI. Using a four part diagnostic framework built from four tests, namely Authorship, Reproduction, Buffer, and Form, the paper demonstrates that both interventions were structured to be legible to funders and evaluators before they were accountable to the people they claimed to serve. Plan 2020 named its own accountability horizon in its title and then produced no public reckoning when that horizon arrived and the disparities it promised to close had instead widened. The AI adoption report performs the same maneuver at a larger scale. It identifies the most relationship dependent work in the nonprofit sector, including psychiatric case management, crisis referral, and therapeutic coaching, as the primary frontier for AI substitution, and it reports only throughput metrics, such as hours saved, reduction in length of stay, and efficiency gained, with no mechanism for verifying whether the recovered capacity returns to relationship or is instead captured as budgetary surplus. The paper concludes that artificial intelligence has not introduced a new failure into civic and nonprofit infrastructure. It has automated an existing failure and done so more quickly and at a greater remove from the people whose care is being restructured.
Keywords: nonprofit management, philanthropic infrastructure, artificial intelligence, program evaluation, dependency theory, social work, civic governance
1. Introduction
Every decade produces a new vocabulary for scaling social good, and every vocabulary arrives carrying the same premise, largely unexamined: that the limiting factor in social impact is capacity, and that capacity is a technical problem awaiting a technical solution. Geographic replication, methods built around training the trainer, and philanthropy organized around outcomes each made this claim in turn (Bridgespan Group, 2014; Candid and GEO, 2011). Artificial intelligence is the newest entrant, and it arrives with a research literature already prepared to certify it (Project Evident, 2026).
This paper takes up a narrower and more falsifiable question than whether artificial intelligence can scale nonprofit impact. It asks whether the institutional apparatus that produces and disseminates claims about the benefit of artificial intelligence to the social sector exhibits the same structural pattern found in an earlier civic planning intervention in the same city, one that involved no comparable technology at all: Indianapolis's Plan 2020. The comparison is not rhetorical. It is diagnostic. If the same four structural failures recur across a purely administrative planning process and a technologically mediated program delivery report, the failure cannot be attributed to the technology, or to the planning methodology. It must be attributed to the institutional position from which both were authored.
The argument proceeds in six parts following this introduction. Section two establishes the theoretical foundation and defines the four tests, namely Authorship, Reproduction, Buffer, and Form, as a diagnostic instrument. Section three applies the framework to Plan 2020 as the specimen from the era before artificial intelligence. Section four applies the same framework to Scaling Impact with AI as the specimen from the current era. Section five discusses what the continuity between the two specimens reveals about the durability of the underlying mechanism. Section six considers what this continuity implies for practitioners, funders, and management scholars evaluating claims about artificial intelligence adoption in the social sector. Section seven concludes.
2. Theoretical Framework: The Four Tests
Philip Selznick's (1949) study of cooptation inside the Tennessee Valley Authority showed that an institution can absorb representatives of an affected community into its own formal structure for the specific purpose of neutralizing their capacity for independent action. This is participation without authorship, and it remains one of the most durable findings in the sociology of formal organization. Albert Bandura's (1999) theory of moral disengagement identifies the cognitive mechanisms, including diffusion of responsibility, euphemistic labeling, and displacement of agency onto systems, through which institutions and individuals distance themselves from the human costs of their own decisions. Ivan Illich, writing with Irving Zola, John McKnight, Jonathan Caplan, and Harley Shaiken (1977), described what they called the disabling professions: expert systems that define a population's needs in terms only the expert system itself is equipped to meet, thereby manufacturing the very dependency the system claims to remedy. John Kretzmann and John McKnight (1993) supplied the positive counter model in their framework for asset based community development, in which capacity building begins with what a community already possesses rather than with what an external institution proposes to supply.
From these four foundations it is possible to derive four tests, and to apply them to any intervention that claims to build capacity, deliver care, or scale impact on behalf of a population it does not itself constitute. Each test asks a different question about where authority, benefit, and accountability actually reside inside an intervention, regardless of whether that intervention is a planning document, a piece of software, or both.
2.1 Authorship
The first test asks who wrote the account of what happened, and in whose voice the intervention speaks. An intervention fails the Authorship test when the people it serves appear in its own record only as data points, as quotations selected by someone else, or as beneficiaries of a tool's action rather than as agents of their own decisions. Selznick's cooptation finding is the clearest precedent for this test: representation inside a structure is not the same thing as authorship of the account that structure produces about itself (Selznick, 1949).
This test is not satisfied by the presence of testimonials, case studies, or quoted beneficiaries. A case study in which a client's outcome is reported by the organization that served them, without any independent account from the client and without a clear description of how consent for their story was obtained, still fails the test, because the organization retains authorship over the narrative even while appearing to include the client within it. What matters is not whether a served person appears in the text, but who controls what is said about them, and why.
2.2 Reproduction
The second test asks whether the intervention builds the capacity of the population it serves to act independently of the intervention, or whether it instead reproduces that population's dependency on the intervention itself and on the institutions that fund and evaluate it. An intervention fails the Reproduction test when its own recommended next step is simply more of itself: more engagement with the same vendor, consultancy, or funder, rather than a transfer of capacity the population could carry forward without that institution's continued involvement.
Kretzmann and McKnight's (1993) asset based framework supplies the positive standard against which this test is measured. A genuinely capacity building intervention should, over time, make its own continued presence less necessary rather than more. Illich and his coauthors (1977) identified the opposite tendency in professionalized systems of care: the expert system defines the need in terms that only the expert system can meet, so that need and remedy expand together rather than the remedy retiring the need over time. The Reproduction test asks, concretely, whether an intervention's own literature functions as an argument for continued dependency on the intervention, however the surrounding language is framed.
2.3 Buffer
The third test asks what happens to the resources, time, or capacity that the intervention claims to free. An intervention fails the Buffer test when it reports savings, efficiency gains, or capacity liberation without any mechanism for verifying whether the freed resource returns to the relationship it was extracted from, or is instead captured elsewhere as budgetary surplus, reduced staffing, or a smaller grant in the following funding cycle.
Bandura's (1999) account of moral disengagement is directly relevant here, because a Buffer failure is rarely a lie. It is more often an omission enabled by euphemistic labeling: a reduction in length of stay, a gain in efficiency, an increase in items processed, each phrased as an unambiguous good, none paired with an account of where the recovered value actually went. The Buffer test does not require proof that recovered capacity was captured rather than reinvested. It requires only that the intervention supply a mechanism by which either outcome could be verified. The absence of such a mechanism is itself diagnostic, and it is the failure this paper documents most consistently across both specimens examined below.
2.4 Form
The fourth test asks whether the structure used to describe, categorize, or evaluate the intervention emerged from the practice itself, or was imported from an external taxonomy and then retrofitted with examples drawn from the practice. An intervention fails the Form test when its own methodology admits, even in passing, that its categories were designed for a different context and only afterward found to correlate with the population under study.
This test matters because form is not neutral packaging. A taxonomy built for one purpose carries assumptions about what counts as a meaningful unit of analysis, and those assumptions travel with the taxonomy into its new application whether or not they actually fit. A specimen that fails all four tests exhibits what this paper terms institutional groundwater contamination, a condition in which the mechanism producing harm is embedded in the very medium the institution draws from and distributes, such that the institution's own remedial language is drawn from the same contaminated source it claims to be purifying.
3. Specimen I: Plan 2020 and the Rhetoric of the Deadline
Plan 2020 was a comprehensive planning process for Indianapolis, co led by the Greater Indianapolis Progress Committee and the Indianapolis Department of Metropolitan Development, conducted between 2014 and 2016 and organized around six pillars: Choose, Connect, Love, Serve, Thrive, and Work Indy. The plan was timed to the city's 2020 bicentennial, a fact embedded directly in its own name. Its advisory structure drew on the same set of institutions that recur throughout the author's broader research into Indianapolis civic governance, including the Central Indiana Community Foundation, the Central Indiana Corporate Partnership, the Indy Chamber, the Local Initiatives Support Corporation, and the city's Health and Hospital Corporation (McAleavey, 2026). The plan received a National Planning Achievement Award from the American Planning Association and reported that its engagement process had reached more than 100,000 residents.
Plan 2020 fails the Authorship test in its process design. Residents were engaged as respondents to a planning apparatus already staffed and funded by the institutions named above, not as coauthors of the plan's substantive commitments. The clearest evidence of this failure is what the plan produced after its own completion: not a resident authored planning capacity, but the People's Planning Academy, a credentialing program that trains residents to understand the comprehensive plan already written on their behalf. Where an asset based model would have residents author the plan and the institution learn to receive it, the People's Planning Academy inverts that relationship, training compliance with a document residents did not write.
The plan's most significant failure, and the one that gives this specimen its evidentiary weight, is the Buffer test, and it becomes visible only in retrospect. Plan 2020 set its own accountability horizon inside its own title. When that horizon arrived, the racial and economic disparities tracked continuously by the city's own civic data infrastructure, including instruments such as the Racial Equity Report Card, had not narrowed (McAleavey, 2026). No public reckoning followed the arrival of the plan's own named deadline. This absence of reckoning is itself the buffer at work: an institution that names its own deadline publicly, and then allows that deadline to pass without any public accounting of whether its commitments were met, has converted what was framed as a commitment into an artifact of public relations, insulated from the very consequence it was originally offered to be judged against.
The Form test is failed by the plan's own structure. Six alliterative pillars, chosen for their communications value and their fit inside a single branded framework, were imposed on a set of civic challenges, including housing, safety, education, and economic mobility, that do not naturally sort into six coordinate categories of equal weight or scope. The form was selected for its rhetorical utility before it was tested against the actual shape of the problems it claimed to organize.
It should be stated plainly what this specimen does and does not establish. The evidence assembled here documents a pattern of process design, institutional composition, and outcome tracking. It does not establish that any individual named above acted in bad faith, and no such claim is made or implied. The argument of this paper concerns structure, not motive.
4. Specimen II: Scaling Impact with AI and the Migration of the Mechanism
Project Evident is a nonprofit consultancy whose OutcomesAI practice provides paid technical assistance to nonprofits and funders on data, evidence, and the adoption of artificial intelligence. In June 2026 the organization published Scaling Impact with AI, the first year output of a multi year study supported by the Siegel Family Endowment. The report compiles publicly available information on 128 nonprofit organizations that use artificial intelligence in program delivery, and it organizes their practices into 18 program model elements, nested under six program categories, and mapped against 13 types of artificial intelligence applications (Project Evident, 2026). Every factual claim about the report's content in this section is drawn directly from the report's own published text and is cited to it accordingly.
4.1 Where the Automation Is Pointed
The report's central empirical finding is that the two leading categories of adoption are Personalized Support and Service Coordination, categories the report characterizes as difficult to scale because the work is driven by relationship and requires individualized attention (Project Evident, 2026). This is not an incidental finding buried in an appendix. It is presented as the report's leading insight, placed first among the four insights highlighted in its executive summary. The frontier of adoption identified by the report's own data is not administrative efficiency, such as fundraising or scheduling, but the relational core of direct practice: psychiatric case management, crisis referral, benefits navigation, and therapeutic coaching.
The report frames this finding as evidence of capacity liberation, freeing staff for higher judgment work. The same finding can be read, with equal fidelity to the data presented, as automation locating precisely the labor that has, until now, resisted industrialization because it depends on sustained human relationship. The report does not adjudicate between these two readings. It states the finding and proceeds directly to case studies that assume the first reading is correct, without engaging the second.
4.2 Authorship
Every case study in the report is narrated from the position of the tool rather than the position of the practitioner or the person receiving care. Gemma Services' therapists do not appear in the report as the agents who used better information to make a better clinical decision. The report instead describes an analytics tool that supplies therapists with real time information used to shape individualized treatment decisions, and it attributes the resulting outcome, a reduction in the length of residential stays, to the tool's function (Project Evident, 2026). Age UK's Telephone Friendship Service callers do not appear as people whose distress a trained volunteer learned to hear. The report describes artificial intelligence being used to transcribe volunteer calls and to identify safety concerns among older callers (Project Evident, 2026).
In each of these cases, and across the eighteen case studies presented in the report's appendix, the practitioner and the person receiving care are present only as the site at which the tool's benefit is measured. Neither appears as an author of the account. This displacement is compounded by a disclosure inside the report's own methods section, which states that artificial intelligence was used to help draft portions of the document itself (Project Evident, 2026). A report that attributes agency over the narration of care to artificial intelligence was itself, by its own admission, partially authored by artificial intelligence. This is not an incidental detail. It is a recursive instance of the exact displacement the report documents as an emerging market pattern, occurring inside the document that is doing the documenting.
4.3 Reproduction
The report's recommendations to practitioners encourage looking toward peer organizations that have already adopted the most widely used of the eighteen program model elements the report itself has taxonomized (Project Evident, 2026). Its recommendations to funders instruct them to build shared infrastructure and curated resource hubs to support continued adoption of artificial intelligence across the sector (Project Evident, 2026). In both cases, the recommended next step is convergence on a framework whose categories were authored by the consultancy issuing the report, paired with continued engagement, explicitly described as a multi year study, with that same consultancy.
The report does not recommend that nonprofits build independent evaluative capacity of their own. It recommends fluency in a taxonomy that only Project Evident is currently positioned to update, interpret, and extend across the remaining two years of its study. This is the Reproduction test applied concretely: an intervention whose own literature functions, whatever its stated intent, as an argument for continued reliance on the institution that produced it.
4.4 Buffer
Every headline metric presented in the report is a throughput metric, and none is paired with an account of where the recovered value went. The report cites a reduction of 29 days in residential length of stay at Gemma Services, a savings of 9,500 staff hours across 23,000 calls at Age UK, 10 hours per week returned to case managers using CaseAI, a 35 percent increase in the number of items a single worker can list at Goodwill Industries of Orange County, and an efficiency gain of roughly 70 percent in crisis referral reported by Signpost AI (Project Evident, 2026). Not one of these figures is accompanied by a corresponding account of whether the recovered time, capacity, or savings were redirected toward deeper relationship with the people served, or captured instead as reduced staffing, reduced future funding, or an increased caseload for each remaining worker.
The report cannot supply this account, and the reason is structural rather than a matter of oversight. Project Evident is a consultancy retained to advise funders and nonprofits on the adoption of artificial intelligence, operating inside a multi year study funded by a single endowment whose stated purpose is to explore the impact of artificial intelligence within the social sector (Project Evident, 2026). The consultancy's continued relevance depends on efficiency being legible as an unqualified good. It does not depend on tracing whether efficiency gains are reinvested in relationship or captured elsewhere. This is the clearest failure of the Buffer test among the four: a report organized entirely around the language of capacity liberation, with no mechanism anywhere in its pages for verifying that liberated capacity is not, in practice, capacity foreclosed.
4.5 Form
The report's own methods section states that its 13 types of artificial intelligence application were developed inductively, meaning derived from observation of the 128 organization dataset rather than imposed upon it (Project Evident, 2026). The same section then discloses that the resulting taxonomy was subsequently compared against two preexisting frameworks not designed for the nonprofit sector: the NIST AI Use Taxonomy, published in 2024, and the OECD Framework for the Classification of AI Systems, published in 2022 (Project Evident, 2026). The report states that nine of its 13 categories map directly or closely onto categories in one or both of these external frameworks.
A claim of inductive discovery and a disclosure of substantial correspondence with frameworks built for general artificial intelligence governance, in contexts entirely outside the nonprofit sector, cannot both be fully true at once. The more economical reading is that a governance oriented taxonomy, built originally for regulatory and industrial contexts, was applied to a set of nonprofit case studies and then confirmed by the degree to which it fit them, in essentially the same manner that Plan 2020's six alliterative pillars were confirmed by their capacity to sort a wide range of loosely related civic domains into a single, communications ready structure. The report's own acknowledgment of this alignment, offered there as a strength, is read here as the clearest documentation available of the Form test's failure.
5. Discussion: The Groundwater Does Not Move, It Finds a New Pipe
The value of comparing Plan 2020 to Scaling Impact with AI is not that the two interventions share a subject matter. One is a municipal planning process. The other is a report on the adoption of a new technology by nonprofit organizations. What they share is a structural position: both were authored by institutions positioned to benefit from the continuation of the problem they claimed to be addressing, and both produced accounts legible to funders and evaluators before those accounts were legible to the people whose lives the intervention claimed to improve.
Ten years separate the two specimens, and across that decade the instrument itself changed, from a comprehensive municipal plan to a machine learning application deployed inside individual nonprofit organizations. The four failures identified by this paper's diagnostic framework did not change. Authorship remained located outside the population being served. Reproduction remained oriented toward continued engagement with the authoring institution rather than toward the transfer of independent capacity. The Buffer remained unaccounted, converting freed resources into a good that the reporting institution asked its audience to accept without verification. Form remained imported from elsewhere rather than derived from the practice under study.
What artificial intelligence appears to have changed is velocity and distance rather than the underlying mechanism. A municipal planning process takes years to fail, and its failure becomes visible only when its own named deadline finally arrives, often a decade later, as happened with Plan 2020. A software deployment, by contrast, can scale from 18,000 users to more than one million users within roughly two years, as the report itself documents in the case of Lenny Learning (Project Evident, 2026). This means the interval between adoption and any possible reckoning with a Buffer failure is compressed by an order of magnitude relative to a municipal plan, while the distance between the person receiving care and the institution authoring the account of that care is simultaneously extended, because the intervening layer is now a proprietary tool rather than a public planning document that residents could, in principle, read for themselves.
This is the sense in which the title of this paper is meant literally rather than as a figure of speech. The same institutions that diagnose civic and philanthropic dysfunction, that fund research into equity, and that publish reports on the promise of artificial intelligence for the social sector, are frequently the institutions whose own prior instruments produced the conditions those later reports are now called upon to address. They teach the diagnostic vocabulary. The evidence assembled in this paper indicates that they have not yet, in either specimen examined here, submitted their own instruments to that same vocabulary.
6. Implications for Management and Philanthropic Practice
For funders evaluating proposals related to the adoption of artificial intelligence in the social sector, the four tests offer a due diligence instrument that does not require technical expertise in artificial intelligence itself. A funder can ask, of any proposal, four plain questions. Who authors the account of the intervention's success, and in whose voice is that account written. What is the intervention's own recommended next step, and does that step lead toward capacity the grantee organization can carry forward independently, or toward continued reliance on the vendor, consultancy, or funder who authored the proposal. Where does freed capacity actually go, and does the proposal include any verification mechanism beyond a reported metric. Did the categories used to describe and evaluate the intervention emerge from the practice itself, or were they imported from elsewhere and then confirmed by selective fit.
For nonprofit management scholars, the persistence of these four failures across one purely administrative intervention and one technologically mediated intervention, separated by a decade, suggests that research on artificial intelligence adoption in the social sector should be evaluated as a continuation of the existing literature on philanthropic governance, rather than treated as a novel category requiring an entirely new evaluative vocabulary. The risk this paper identifies is not that artificial intelligence introduces an unprecedented harm into program delivery. The risk is that research on artificial intelligence adoption inherits, without examination, the same structural blind spots that have characterized capacity building interventions in this sector for at least a decade, and does so at a scale and a speed that make the kind of retrospective accounting performed in this paper, applied to Plan 2020 a full decade after its own named deadline, considerably harder to perform in anything close to real time.
Practitioners inside nonprofit organizations adopting these tools are not without recourse. The four tests can be applied internally, before a technology proposal from a vendor or consultancy is accepted, by asking the same four questions a funder would ask, and by insisting that any reported efficiency gain be paired with a public, written account of where the recovered capacity was redirected. An organization that cannot answer where its own freed hours went has failed the Buffer test regardless of whether the technology it adopted performed as advertised.
None of this implies that artificial intelligence itself is unsuited to nonprofit program delivery. The dataset assembled by Project Evident documents real reductions in length of stay, real hours returned to case managers, and real increases in the number of people an organization can reach. The argument of this paper is narrower, and for that reason more durable: these real gains are being reported inside a structure that has not yet been asked to account for itself by the same four tests it would, in another era, have been the first to apply to someone else's plan.
7. Conclusion
Plan 2020 named its own deadline and then let that deadline pass without a public accounting. Scaling Impact with AI reports its metrics without naming any deadline for an accounting at all. Between the two lies a full decade of continuity: the same four failures, the same institutional position, the same displacement of the person receiving care into a data point inside someone else's evidence of success.
The question this paper leaves for management scholars, for funders, and for practitioners is not whether artificial intelligence can scale nonprofit impact. The dataset assembled by Project Evident suggests, descriptively, that it already has. The real question is whether the social sector will submit that scaling to the same four tests it has, so far, failed to apply to its own prior instruments of care, or whether it will continue teaching the vocabulary of contamination while pouring, at increasing speed, from the same pipe.
References
Bandura, A. (1999). Moral disengagement in the perpetration of inhumanities. Personality and Social Psychology Review, 3(3), 193 209.
Bridgespan Group. (2014). Transformative scale: The future of growing what works. The Bridgespan Group.
Candid and GEO. (2011). What do we mean by scale? Grantmakers for Effective Organizations.
Illich, I., Zola, I. K., McKnight, J., Caplan, J., and Shaiken, H. (1977). Disabling professions. Marion Boyars.
Kretzmann, J. P., and McKnight, J. L. (1993). Building communities from the inside out: A path toward finding and mobilizing a community's assets. ACTA Publications.
McAleavey, M. (2026). Quality of Liberation Planning: A comprehensive repair strategy built by failure. Unpublished manuscript. Joy Repair.
Project Evident. (2026). Scaling impact with AI: Emerging patterns in nonprofit program delivery. Project Evident, with support from the Siegel Family Endowment.
Selznick, P. (1949). TVA and the grass roots: A study in the sociology of formal organization. University of California Press.
Note on evidentiary basis: all claims regarding the content of Scaling Impact with AI in section four are drawn directly from that report's own published text as retrieved on July 15, 2026, and are cited to it throughout. Claims regarding Plan 2020 in section three draw on the author's prior documented research into Indianapolis civic governance and are cited to that manuscript. This paper makes no claim regarding the intent or state of mind of any individual named in either specimen; its argument is structural and does not extend to motive.

