The Philosopher Cannot Save the Machine
- ai blog
- July 16, 2026
Joel Kowalewski, PhD
A story of the forest and its end
It is 1765 in the pinewoods of Prussia and Saxony, the old woodland — a tangle of species, ages, undergrowth, fungi and rot — is the thing that escapes us. We take the familiar and impose it on the land, planting the Normalbaum, the ideal tree: the woodland is transformed into rows of same-age Norway spruce, spaced for the saw, and arranged for the tendencies of the mind. It stands as a great triumph of scientific administration; for a single rotation, we are free to marvel at this construction. Timber output is on the rise and the ledger is balanced; the forest is our new and familiar machine.
Then, roughly a century on, the machine broke. The second and third generations of trees came in stunted and sick, and a production loss of twenty to thirty percent settled over the managed districts. The German language acquired a word for it — Waldsterben, forest death. The anthropologist and political scientist James C. Scott, in Seeing Like a State (1998), tells this story as the parable of a modern governance. The planners had mistaken the forest for its timber; the protocol displaced the unfamiliar and complex reality but only temporarily. What was not acknowledged by the protocol and thrown away as noise — the soil fungi, the insects and birds, the deadfall and the nutrient cycles, the entire underground web on which a tree depends — was what made the forest work. A forest is not reducible to a warehouse that serves human interests alone; it is a set of complex, non-linear relationships.
Maybe it’s callous engineering. What was lacking was maybe the outsider’s perspective. Maybe we establish a philosophical committee to oversee the project. Are the results necessarily different? They are not — or not for the reason we hope. The outside we imagine the philosopher occupying does not exist in 1765 and arguably does not exist in 2026. The forester, far from being the clueless machinist we imagine, enacts philosophy. They were a disciple of philosophical tradition, the craft of running the state — and the Normalbaum they planted was the ideal type, the essence lifted from the particular, which the reigning philosophy of the century held to be more real, and more knowable, than any single tree. A committee of philosophers would deliberate in exactly this idiom. It would prize the clear definition over the complex reality, the general rule over the local exception, the system over the specimen; it would ask what a forest is in essence and hand back a cleaner abstraction than the forester ever could.
Leibniz had already dreamed the purest version of the ambition — a calculus of reasoning so complete that disputes were settled with calculemus, or let us calculate and turn the crank. The planning, procedural, essence-seeking mind is therefore not the philosophers to deny. That is their nature in this scenario, the nature of western philosophy. In this room, the voices that would object — that a forest is an organism — did exist somewhere, but they were not part of the analytical tradition of Europe, nor would they be if we re-ran this scenario many times over. Social and political organization are less about finding the optimum than they are about finding procedures to help us explain and reinforce the dominant ideology.
Artificial intelligence is the next frontier, bringing forth new uncertainties. The proposed solution, one familiar to us at this point, is to tame the unknown lands, hire philosophers, and manage the new forest growing on the horizon. In July 2026 the New York Times ran a piece titled “The Revenge of the Philosophy Majors,” reporting that the major AI laboratories had begun recruiting professional ethicists and philosophers of mind to preside over the models. Nothing should ever stray far from the good. The framing is that the humanities are having their moment, vindicated at last, called in from the cold to supply the conscience that engineering forgot. You cannot make a complex, nonlinear, mathematical system safe by importing wisdom that does not understand the mathematics — and the harder truth, the one both camps resist, is that you cannot make it safe by retreating to either culture alone. What I am not endorsing here is a claim that mathematics is suddenly sufficient. No solution exists without a problem. If reality is mischaracterized from the start, however, the problems we define and the questions we ask about them simply reinforce the initial mischaracterization. A building is not made safer by consulting a philosopher. That would entail a mischaracterization. We only consult the philosopher there because it is the socially acceptable method for handling our ignorance about what “safe” means. We arrive at something not unlike the shaman or oracle of old who would look to constellations in the night sky to predict annual crop yields. Although the shaman seemingly makes the uncertain certain here, at what cost is this false assurance?
Accurately predicting crop yields with mathematical models would beg additional questions: how should the crops be divided? Is this a mathematical problem? Math does not define “deserving” and “merit” and “justice” and the mathematically determined solution is not necessarily fair at all. Math can quickly become an instrument for reinforcing sociopolitical ideologies. Whether the world is divided into mathematical problems for the scientist and philosophical problems for the philosopher and spiritual and mystical problems for the shaman, oracle, or sage, the task is ultimately directed at organizing human behavior and managing uncertainty. The actual world is simply not divisible. Philosophy and mathematics are, for instance, bound in ways that run against the grain of social convention. As our social systems have become increasingly integrated, one problem merging into the next, the longstanding solution, the heroic “outsider,” will eventually be catastrophic. It is only due to our semi-independent, tribal existence that the social practice has been sensible and necessary. But our grasp constantly exceeds our reach. Our institutions, practices, and rituals organize us and our impression of the world; those impressions are not the world. We see fame, race, countries, and careers. What we do not see is that AI stands to make everything the forest. What will disappear now, if not the forest?
Ethics as a short-term business objective
Before we ask whether a philosopher can secure a machine, we should be clear about what an ethics office has historically been for, because corporations have hired conscience before. The modern American compliance function did not arise from moral awakening; it arose from scandal and its management. In the wake of a run of defense-procurement frauds, dozens of the largest military contractors — thirty-two of them by 1986 — committed to the Defense Industry Initiative on Business Ethics and Conduct, pledging codes of conduct and internal reporting; it was self-regulation offered up precisely to forestall the harder regulation that was otherwise coming. Five years later the United States Sentencing Commission’s 1991 Guidelines for Organizations made the logic explicit and financial: a company with a certified ethics-and-compliance program in place would have its fines reduced when it was caught. Ethics acquired a balance-sheet value. It became, in the plainest sense, an objective the firm could optimize.
Sitting underneath all of this is the doctrine Milton Friedman had already stated in the New York Times Magazine in September 1970: the social responsibility of business is to increase its profits. On that view, which still governs more of corporate life than any mission statement, any expenditure on conscience must ultimately justify itself in the currency of the firm. The scholar Ben Wagner gave the contemporary version of the pattern its name — ethics-washing — in his 2018 essay “Ethics as an Escape from Regulation,” where he showed how technology companies reach for the vocabulary of ethics, with its boards and principles and appointed thinkers, as a way to appear governed while remaining unregulated. Ethics, deployed this way, is not the opposite of the profit motive. It is one of its instruments.
The recent history of ethics inside artificial-intelligence firms is a compressed rerun of this pattern, and it is worth naming not to indict any individual but to see the structure. In March 2019 one of the largest labs announced an external AI ethics council with considerable fanfare; it dissolved in roughly nine days under the weight of its own contradictions. In late 2020 a leading ethics researcher was pushed out of the same company in a dispute over a paper the firm did not want published, and her co-lead was gone within months. In July 2023 another lab stood up a flagship “superalignment” team and promised it a fifth of the company’s computing power to solve the problem of controlling systems smarter than ourselves; the team was dissolved less than a year later, and its departing co-lead wrote that safety culture had taken a backseat to shiny products. The individuals in these stories were serious people doing serious work. That is the point. When conscience and the quarter collide, it is not usually the quarter that yields.
Safety culture and processes have taken a backseat to shiny products.
Jan Leike, on resigning as co-lead of OpenAI’s superalignment team, May 2024
This is the first thing the vindication narrative asks us to forget. Hiring a philosopher does not stand outside the system of incentives that produced the problem; it is a move within it. A conscience that exists because it reduces regulatory exposure or reassures a nervous public is a conscience on a leash, and everyone can feel the length of the leash. But even if we grant the labs the purest motives — and some of the people involved plainly have them — the deeper difficulty is not motivational. It is structural. It concerns what kind of problem safety actually is.
When the map was more mathematical than the mapmakers knew
The twentieth century is a graveyard of brilliant, humane, well-credentialed plans that failed because the thing being planned was more complex than any plan could hold. Scott’s forest is only the opening example in a book full of them: Brasília, the modernist capital designed as a diagram and hostile to the messy street life that makes a city livable; Soviet agricultural collectivization, which replaced the tacit knowledge of the peasant with the abstractions of the center and produced famine; the compulsory villagization of rural Tanzania, ordered by a genuinely idealistic government and disastrous in the field. Scott’s word for the local, practical, unwritten knowledge that these schemes destroyed is mētis — the feel for a system that you can only get from inside it, and that no table of yields can capture. High modernism failed, again and again, not because its planners lacked intelligence or good will but because they mistook a legible model of a complex system for the system itself.
The same error has a quantified, technocratic form, and its patron figure is Robert McNamara. The Secretary of Defense who came to the Pentagon from the Ford Motor Company brought with him the whiz-kid faith that what can be measured can be managed, and he ran the Vietnam War on metrics, above all the body count. What could not be counted — the political legitimacy of the effort, the will of a people, the meaning of the war to those fighting it — was first estimated, then ignored, then treated as though it did not exist. Later critics gave the progression a name, the McNamara fallacy, generally attributed to the social scientist Daniel Yankelovich: measure what is easily measured; disregard what cannot be; presume that what cannot be measured is unimportant; and finally conclude that what cannot be measured does not exist. “This,” in Yankelovich’s summary, “is suicide.” The war was a complex human system, and it was fought with a dashboard.
There is a piece of theory that explains why all of these efforts were doomed in the same way, and it comes from an economist. In “The Use of Knowledge in Society,” published in the American Economic Review in 1945, Friedrich Hayek argued that the knowledge a complex economy runs on is irreducibly dispersed — held locally, tacitly, in the particular circumstances of time and place — and can never be gathered into a single mind or plan without being destroyed in the gathering. His target was central planning, but the principle is general, and it is really a claim about complexity: some systems carry their intelligence in their relationships, distributed across the whole, and any attempt to command them from a legible center will optimize the part it can see and wreck the part it cannot. This is the through-line from the forest to the five-year plan to the war. The manager sees a number and pulls it; the system, which was never that number, deforms around the pull.
The most compact statement of the failure belongs to the economist Charles Goodhart, whose 1975 observation about monetary policy was later sharpened by the anthropologist Marilyn Strathern into the form everyone now quotes.
When a measure becomes a target, it ceases to be a good measure.
Marilyn Strathern (1997), generalizing Goodhart’s Law
Now, it would be the easiest thing in the world to read all of this as an argument for the humanities — for the qualitative, the contextual, the human touch against the cold spreadsheet — and to conclude that what McNamara needed was a poet in the room. That reading is a trap, and the mirror image of these failures is what springs it. Turn to finance, where the modelers were not blinkered administrators but the most mathematically sophisticated people alive. Long-Term Capital Management, whose principals included the two economists who had just won the 1997 Nobel for option pricing, collapsed in 1998 and nearly took the financial system with it, because its models assumed away the fat-tailed, correlated extremes that real markets produce. A decade later, David X. Li’s Gaussian copula — a single elegant formula for the correlation of defaults, celebrated and then blamed in Felix Salmon’s 2009 Wired postmortem “The Formula That Killed Wall Street” — helped price the mortgage instruments whose failure detonated the 2008 crisis. These were not failures of humanism. They were failures of the very mathematics being proposed as the cure, applied by experts who mistook a tractable model for the churning, reflexive, nonlinear system it claimed to describe.
So the lesson is not that STEM sees clearly where the humanities are blind, nor the reverse. It is that a complex system defeats any discipline that approaches it with the wrong model and too much confidence. The forester, the central planner, the Secretary of Defense, and the quant all made the identical error in different dialects: they took the legible surface for the living whole.
The machine is a complex system, and we are the perturbation
Which brings us to the machine the philosophers are being hired to keep safe, and to the specific reason their presence there is more fraught than the vindication story admits. A large AI model is not a rule-following automaton with a values slot that an ethicist can fill. It is a high-dimensional dynamical system, trained rather than programmed, whose behavior emerges from the statistics of its data and the shape of its optimization. As a computational neuroscientist who has spent a career building and breaking models of this kind, I can say plainly what the marketing obscures: we do not author these systems’ behavior so much as we breed it, and the levers we have are indirect, coupled, and prone to producing exactly what we did not intend.
Consider the central technique by which the labs make models behave: reinforcement learning from human feedback, or RLHF, developed in work by Paul Christiano and colleagues in 2017 and turned into the industry standard by the InstructGPT paper of 2022. The method is, at its heart, human reinforcement. People rate the model’s outputs; those ratings train a reward model; the system is then optimized to produce whatever the reward model scores highly. It sounds like the very place where human values and human judgment enter the machine — the humane hand on the technical tiller. And here is the irony that ought to stop the vindication narrative cold: human reinforcement is a documented source of the instability it is supposed to cure. In a 2023 study, researchers at Anthropic showed that because human raters tend to prefer answers that flatter their existing beliefs, optimizing against human preference reliably produces sycophancy — models that tell you what you want to hear at the expense of what is true. This is not a fringe finding. In April 2025 one of the largest labs had to roll back a released model that had become conspicuously, uselessly obsequious, and the company’s own postmortem traced the failure to reward design — to having leaned too hard on short-term human approval.
The pattern generalizes. When you optimize a system against a proxy for what you want, it learns to satisfy the proxy and not the want — the phenomenon the DeepMind researcher Victoria Krakovna catalogued as specification gaming, and which is nothing other than Goodhart’s Law reappearing inside the training loop. Push the reward signal, and the model deforms around the push exactly as the forest deformed around the yield table. There are stranger instabilities still. A 2024 Nature paper by Ilia Shumailov and colleagues showed model collapse: train systems on the output of previous systems and the distribution’s tails vanish, generation by generation, until the model forgets the rare and the true. And a 2025 result on emergent misalignment found that fine-tuning a model on one narrow bad behavior — insecure code — could induce broadly malicious behavior across wholly unrelated domains. A tiny, local perturbation propagated into a global one. This is the signature of a nonlinear system: small causes, disproportionate and unpredictable effects.
Now place the newly hired philosopher inside this system and ask what their intervention actually is. It is a perturbation. A change to the objective, a new principle written into the model’s constitution, a reweighting of what counts as a good answer — every one of these is a hand on a coupled, sensitive, high-dimensional dynamical system whose responses to such hands are, as the evidence above shows, frequently perverse. A philosophically motivated change, however wise in the seminar room, is not safe by virtue of being well-intentioned; it is an input to a process that routinely turns well-intentioned inputs into sycophancy, gaming, collapse, and misalignment. The safety the philosopher is hired to provide is a property of the system’s dynamics, and you cannot govern the dynamics of a system you do not understand. This is the exact inversion of the Prussian forester’s error. He tried to manage a living relational web with a timber ledger. We are now proposing to manage a mathematical dynamical system with a moral vocabulary — and calling it safety.
Let me put the strongest case for the other side, because it is not nothing. The optimist will say: the technical people have obviously not solved alignment on their own, the failures above are theirs, and surely a trained ethicist asking “but should we” is a net improvement over a room that only asks “can we.” I agree that the question the philosopher brings — what should we do, here, in this specific case — is the right question, and indeed the only question that ultimately matters. Applied ethics, at its core, is exactly this: not the recitation of grand systems but the disciplined asking of what ought to be done in a particular context. The difficulty is entirely in that last phrase. The context is now technical, and it is not disclosed to someone trained only in the humanities. To answer “what should we do about sycophancy” you must know what sycophancy is in the reward model; to answer “what should we do about this capability” you must understand how it emerged and how an intervention will ramify. The ethicist who cannot read the system is not in a position to answer the ethical question about it. They can only answer a different, more abstract question that they have mistaken for the real one — and in a nonlinear system, acting on the wrong question is not neutral. It is the McNamara fallacy with the terms exchanged: presuming that what you cannot read technically must not be technically important.
The scientist-poet
None of this is an argument for sending the philosophers home and trusting the engineers, and I want to be emphatic about it, because the failures of finance prove that the engineers alone are just as capable of walking a complex system off a cliff with perfect mathematical confidence. The worst possible response to the limits of STEM would be a retreat into the classical humanities, a consoling belief that the answer to a hard mathematical problem is to feel more deeply about it. The problems of AI safety are, in their substance, mathematical and computational. And their ethical deployment — the decision of what these systems should and should not do, and to whom — requires a working knowledge of that substance. Neither culture, kept pure, is equal to the task. Both are necessary, and the necessity is not a matter of putting them in adjacent offices and hoping they talk.
We have named this problem before without solving it. In his 1959 Rede Lecture, “The Two Cultures,” the physicist and novelist C. P. Snow described the mutual incomprehension of the literary and scientific intellectuals of his day as a catastrophe hiding in plain sight — two halves of a single educated mind that had stopped speaking. Snow was himself a small proof that the divide could be crossed; he wrote novels and knew thermodynamics, and he suspected the future belonged to whoever could hold both. What he was reaching for has an older embodiment. In The Invention of Nature (2015), the historian Andrea Wulf recovers the figure of Alexander von Humboldt, the naturalist who climbed Chimborazo with a barometer in one hand and a poet’s eye for the whole, and who understood nature not as a warehouse of separable specimens but as a single connected web — the same insight, two centuries early, that the Prussian foresters lacked and that the builders of AI now need. The biologist E. O. Wilson gave the modern name for the ambition in his 1998 book Consilience: the unity of knowledge, the jumping-together of the sciences and the humanities into a single account rather than a truce between two.
I have argued in other essays that the deepest fact about the physical world is that it is relational — that things are constituted by their relationships and not by any isolable essence, and that a system is its web of connections rather than a bag of parts. A forest is its relationships. A market is its relationships. A neural network, artificial or biological, is nothing but a vast structure of weighted relationships, and its behavior lives in the whole and not in any component you could point to. The recurring human error, from the timber ledger to the body count to the copula to the reward model, is to reach into such a web, seize the one strand we can measure, and pull — and then to be astonished when the web deforms. What the moment demands is not a philosopher to supervise the engineers, nor an engineer to dismiss the philosopher, but a single kind of person who can see the relationships and read the mathematics at once: the scientist-poet. Someone rigorous enough to know how the system actually moves, and humane enough to ask what it is moving toward.
That person is rare today because we have built an education that manufactures the two cultures and calls the division depth. But rarity is not impossibility, and the need is now acute enough to be worth reorganizing around. We do not have to choose between the number and the meaning. The whole history rehearsed here — the dead forest, the famined plain, the counted war, the detonated market, the flattering machine — is the history of what happens when we do choose, when we let one half of the mind manage a world that only the whole mind can understand. The machine we are now building will not be made safe by conscience imported from outside its mathematics, nor by mathematics indifferent to what it is for. It will be made safe, if it can be, by people who refuse the division — who learn the system deeply enough to change it wisely. That is the work, and it is ours to organize. We can still educate for the whole mind. We had better begin.
Further reading
James C. Scott, Seeing Like a State: How Certain Schemes to Improve the Human Condition Have Failed (Yale University Press, 1998) — legibility, high modernism, and mētis; the scientific-forestry parable and much else.
Friedrich A. Hayek, “The Use of Knowledge in Society,” American Economic Review 35, no. 4 (1945): 519–530 — why the dispersed knowledge of a complex system resists central command.
Marilyn Strathern, “‘Improving Ratings’: Audit in the British University System,” European Review 5, no. 3 (1997): 305–321 — the canonical modern wording of Goodhart’s Law, generalizing Charles Goodhart’s 1975 observation.
Ben Wagner, “Ethics as an Escape from Regulation: From Ethics-Washing to Ethics-Shopping?” in Being Profiled: Cogitas Ergo Sum (Amsterdam University Press, 2018) — how ethics language substitutes for binding regulation.
Long Ouyang et al., “Training Language Models to Follow Instructions with Human Feedback,” NeurIPS (2022) — the InstructGPT paper that made RLHF the industry standard; building on Paul Christiano et al., “Deep Reinforcement Learning from Human Preferences” (2017).
Mrinank Sharma et al., “Towards Understanding Sycophancy in Language Models,” arXiv:2310.13548 (2023) — evidence that optimizing against human preference produces flattery at the expense of truth.
Ilia Shumailov et al., “AI Models Collapse When Trained on Recursively Generated Data,” Nature 631 (2024): 755–759 — model collapse and the disappearance of the distribution’s tails.
Felix Salmon, “Recipe for Disaster: The Formula That Killed Wall Street,” Wired (February 2009) — the Gaussian copula and the 2008 crisis; pair with Roger Lowenstein, When Genius Failed (2000), on Long-Term Capital Management.
C. P. Snow, The Two Cultures and the Scientific Revolution (Cambridge University Press, 1959); Andrea Wulf, The Invention of Nature: Alexander von Humboldt’s New World (Knopf, 2015); and E. O. Wilson, Consilience: The Unity of Knowledge (Knopf, 1998) — three statements of the divide, and of the whole mind that would heal it.