Can Musicians Achieve Irreplacability? An ALife view
Below is an essay on the question of what might performing musicians have that cannot be replaced by technology – not only what we intrinsically might be able to provide for society, but what might be remunerable. I draw heavily on concepts of artificial life and a previous essay by Angus Lee, "Playing: Or, the subversive politics of interpretation in music" listed in the references.
1. Adopting a Question from Biomedical Engineering
In his recent article "Artificial intelligences: A bridge toward diverse intelligence and humanity's future" Dr. Michael Levin, bio-engineer and researcher of artificial life, poses the question: "We don't want to be replaced, but what are we?" (Levin, 2025). This question resonates with many, and my musician-self started ruminating on our collective existential crisis due to threatened technology-related income loss. Of course, a general de-funding of art and music is happening and is a lower-hanging anxiety to grab on to. Yet, I think there is a connection; if governments, corporations, and other funding bodies witness the ease in which musicians (and much artistic work) can be cheaply replaced - they may ask themselves: "why fund them? Why invest in arts education? It's better to invest in the money-saving technology itself!"
So in response to Levin's question ("what are we?"), I wondered if there is a specific "what" that we performing musicians posses that can't be easily replaced. Levin poses this question to humanity in general in his fascinating essay about synthetic biology, bioengineered bodies, and the coming proliferation of minds that do not fit our existing categories of natural and artificial. His argument, in brief, is that the anxious debate over whether artificial intelligence will "replace" us rests on a mistaken premise: that we are a fixed, bounded, material thing that could be swapped out for a competitor. Levin proposes instead that we are already "an extended, flexible, and adaptive work in progress" (Levin, 2025, p. 4), and the more productive question is not "what we are made of" but what kind of ongoing relationship ("interaction protocol") we are prepared to have with the diversity of bodies and minds already among us (Levin, 2025). This proposal, important as it is for those we deem "others", for bioethics and for thinking about cyborgs, cell colonies, and chimeric organisms, does not translate cleanly into the practicalities of a professional musician's worry about income loss, and I will not try to force it here (but in my imagination, yes). For now, I'll focus on the following emphasis: we don't want to be replaced, but what are we?
Here I will focus on live musical performance and interpretation, asking what, if anything, an individual musician or a collective ensemble possesses that a sufficiently good machine or AI system cannot straightforwardly take away. I don't have an answer. But I will share some thoughts based on my readings, especially around my current research field of artificial life (Alife).
One serious and often-recommended response to replacement anxiety is for musicians to become technologically fluent themselves: to learn enough about AI, machine learning, sound synthesis, and interactive systems to shape the tools rather than merely suffer them, or at least to have a vocabulary with which to negotiate with the people who build those tools. I am a strong advocate of techno-literacy for musicians, but this is not what I am going to talk about. The question here is more pointed: barring the option of becoming a system-builder oneself, what musical qualities can performing musicians rely on to resist replacement?
2. Three Things We Probably Can't Rely On
2.1 Exactitude
I think this is a no-brainer, but there are some interesting things to consider. I remember Brian Ferneyhough once quipped in a masterclass that only the 20th century has equated exactitude with interpretation. Despite the extraordinary complexity and detail of his notation, even he didn't expect it.
Angus Lee, flutist, conductor, and composer makes the case against exactitude in some detail in an article for Vantage. For a machine, he argues, musical notation is information to be counted: symbols are "translated… into numeric instructions as part of an operational sequence," and the resulting performance is qualified chiefly by its "exactitude and velocity" (Lee, 2024, p. 22). Because a machine's relationship to the score is numerical rather than interpretive, it can achieve a kind of faithfulness that would cost a human performer years of tedious repetition, if it is attainable at all. Lee's broader point is that this is precisely why exactitude cannot be where musicians make their stand: a contest of pure execution is a contest a calculating system was built to win.
2.2 Human-ness
A more intuitively appealing fallback is simply being human — trading on the idea that listeners want, at some primal level, to know that another person is on the other end of the sound. There have been studies that show this to be true. Ansani et al. (2025) showed identical audio-visual recordings of piano performances to listeners, telling half of them the pianist was playing live and the other half that the same notes were being produced automatically "thanks to an AI." Listeners rated the ostensibly human performance as more likeable, more emotionally engaging, and of higher quality, even though the sound itself never changed (Ansani et al., 2025). The preference, in other words, does not obviously reflect anything in the actual music; it reflects a belief about its source.
That finding is worth taking seriously and should give us pause, for two reasons that Ansani et al.'s own data supply. First, when the researchers asked participants what differences they had actually noticed between the two versions, the participants "confabulated about differences in rhythm, tempo" and other musical parameters that were actually identical in both versions (Ansani et al., 2025, p. 1). The preference for human-ness, then, is not obviously a perceptual discrimination that a listener is defending with good evidence; it looks more like a prior belief generating its own post-hoc justification. Second, the bias was "moderated by… attitudes toward AI" (Ansani et al., 2025, p. 1), which means not based on the human ear but on a culturally shaped response.
2.3 Liveness
A closer version of the human-ness defense narrows the claim to liveness specifically. Does not being co-present with with a human performer make a significant (and remunerable) difference? Neuroscience research suggests that this is significant. Trost et al. (2024) had listeners undergo brain imaging while hearing the same music performed live versus played back as a recording, and found that live music "stimulate[d] the affective brain… more strongly and consistently than recorded music," producing dynamic, real-time couplings between the performer's playing and the listener's brain activity that recordings simply did not reproduce (Trost et al., 2024, p. 1). This is about as close as the current literature comes to a biological argument that liveness matters; it is a distinct neurophysiological event, arising from "the dynamic relationship between performing artists and the audience" (Trost et al., 2024, p. 1).
Liveness, however real its effects, is not, on its own, a wholly defensible economic promise. It can certainly be capitalized on, though, if you perform a popular genre of music. However, there are several things to consider. First, liveness is a relational event, not a property a musician can privately possess and sell; it requires an audience prepared to show up, which caps the market (unless you are a pop mega-star) no matter how compelling the neuroscience is. Second, nothing in Trost et al.'s design rules out that a sufficiently interactive machine performance — one that responds to a room in real time rather than simply replaying fixed audio — could produce some of the same coupling effects. The important variable in their study is the relational dynamics and their contingent unfolding, not whether we were listening to a being that was carbon-based (human) or silicon-based (machine). Third, history shows liveness has not been enough to keep the bulk of performers economically afloat on its own. Even before Spotify and other streaming platforms, the twentieth-century saw recorded and broadcast music eroding the economic focus of live events.
3. Interpretation as Provocation
What about the appeal of interpretation itself? Not as faithful realization of a text, but as its deliberate, visible, and audible distortion unfolding in real time?
Antonin Artaud's Theatre of Cruelty offers one extreme version of this idea (Artaud, 1958). Artaud wanted a theatre that assaulted an audience's senses. A theater that used the performing body, sound, and spectacle to produce an experience no text or recording could substitute for, precisely because the performance's (sometimes violent) flaunting of convention was the point. Transposed into music, a version of this idea would hold that a performer's irreplaceable contribution is risk: the audible, sometimes uncomfortable evidence that a real person is taking chances with the material in front of witnesses, right now, with something at stake.
Lee arrives at something structurally similar from an entirely different route. His close reading of Glenn Gould's 1965 recording of Mozart's Piano Sonata No. 11 brought to mind Artaud's rejection of European logocentrism since both Gould and Artaud (to different degrees, to be sure!) have critical attitudes toward the "text". Lee contrasts several standard recordings of the Sonata's first movement, all of which faithfully observe Mozart's repeats and tempo markings, with Gould's, which omits the repeats entirely and begins the theme at a tempo bordering on adagio despite Mozart's andante marking. This is a choice Gould himself, in a televised interview, called "perverse" (Lee, 2024, p. 27). Lee proposes that this apparent infidelity is where interpretation actually happens: by refusing literal compliance with the score, Gould exposes a structural logic in the variations — their "progressive, structured acquisition of complexity" (Lee, 2024, p. 26) that the more dutiful performances, in their very correctness, fail to reveal.
Lee is careful, though, to distinguish subversion from mere caprice: "[s]ubversion… by no means implies anarchy," and not every work "deserves, or benefits from" Gould-style interventions (Lee, 2024, p. 27). What licenses the risk is a form of critical judgment about when a convention has calcified into something worth resisting. Gould was concerned that a culture obsessively engaged in "documenting its own [artistic] idioms" would produce art that merely redistributes "selected principles… harvested from other eras" (as cited in Lee, 2024, p. 27) rather than genuinely new interpretive possibility. This concern articulated in the 1960s reads uncannily like a description of the archive-trained generative models we have in 2026! Lee's proposed way forward, then, is not to retreat from technology but to insist that cultural sophistication about interpretation keep pace with technological sophistication. (More about what that might possibly mean later.)
I wonder though, if "provocative interpretation" is really off-limits to a machine. A generative AI model could be prompted to omit repeats, invert dynamics, or "interpret against the grain" of a score. Could a sufficiently well-trained AI model with sophisticated prompting not produce apparent Gould-style subversion by performing a calculated risk? Lee does not, in fact, straightforwardly rule this out. However, the "critical conscience" behind Gould's choices (for example, the judgment about which conventions currently deserve resistance) is not something a system merely optimized for surprising output straightforwardly possesses, nor does it posses a motive. (Lee, 2024, p. 27). Whether this is simply a temporarily unclosed gap, is something time will tell.
For Lee, the crux of the whole argument is: a machine-like execution is "teleological," oriented toward the notation as a fixed goal, whereas a human interpretation is "ontological," oriented toward the notation as raw material for the question "what if?" (Lee, 2024, p. 22).
4. Through the Lens of Artificial Life
When I hear "what if" questions, I can't help but make parallels to my current research in artificial life (ALife). Christopher Langton's founding definition of the field urged researchers to study not "life as we know it" but "life as it could be" (Langton, 1989), asking "what if?" as an invitation to treat "livingness" as a space of possible organizations rather than a fixed biological fact. Interestingly, ALife was founded on almost exactly the same conundrum musicians now face: what, if anything, distinguishes a "real" living thing from a startlingly good simulation of one? Does that distinction actually matter for how we treat, evaluate, or respond to the thing in front of us?
Artificial life visual art embraced these questions from the start, producing artworks that used the vocabulary of biology (growth, evolution, reproduction, metabolism) without particular concern for whether the resulting system was, in some deep sense, "really" alive (Whitelaw, 2004).
This is a useful view for my purposes. The interesting question in ALife art was never really "is this actually alive?" in some final, deterministic sense; it was "what relationship am I, the observer, willing to have with something that behaves as if it might be?" Artist Simon Penny, who has spent decades building interactive artworks in this tradition, makes the observer's role explicit, drawing on second-order cybernetics1: "everything said is said by an observer," as Humberto Maturana put it (Maturana, 1980), meaning that whether a system counts as "lively" is not a fact discoverable in the system alone but an attribution made by someone in relationship with it. Penny's own artistic aim was never to build something lifelike, in the sense of resembling an organism, but something lively, in the sense of behaving with the kind of responsive, unpredictable contingency that makes an encounter with it feel like an interaction rather than a playback (Penny, in press). He coined the term "spectactor" for the audience member in this scenario, because the old binary of active performer and passive viewer breaks down once the audience's own behavior is what the system is responding to (Penny, in press).
Musicologist Christopher Small's concept of musicking independently arrives at a version of the same relational insight. For Small, music is not a set of finished objects (scores or recordings) that listeners passively receive; "to music is to take part, in any capacity, in a musical performance" (Small, 1998, p. 9), and the meaning of a musical event lies not in the notes but in the whole web of relationships — between performers, listeners, and even the people who set up the chairs — that the event brings into being (Small, 1998). If Small is right, then a musical performance was never primarily a transmission of acoustic information in the first place; it was always closer to what Penny's spectactor experiences with a responsive artwork. Again, liveness is less a living presence and really about relationship, which is a distinction with real consequences for what a machine would actually need to replicate in order to replace us.
Penny's spectactor systems are, however, technological artifacts themselves — arguments for building better, livelier machines, not arguments that machines cannot eventually be lively in this sense. If liveliness is, per Maturana, an attribution made by an observer rather than an objective property, there is nothing in that framework that permanently prohibits a sufficiently responsive AI performer from earning the same attribution a live human earns. Instead, it only relocates the contest from "is it human" to "does it produce a convincing relationship," which a well-engineered interactive system might, in principle, do rather well. IRCAM's SOMAX 22 or George Lewis's Voyager3 come to mind.
5. Does This Suggest a Way Forward?
Concepts which have been adopted by ALife visual artists seem easily transferable to music and indeed, a handful of composers and performers have done so already. Could these concepts be more widely adopted in aid of Lee's proposed way forward: not to retreat from technology but to promote cultural sophistication so that it keeps pace with technological sophistication? If sophistication in this context means increasing refinement, savoir faire - knowing how something is done, perhaps so. At least it will give us a vocabulary with which we can discuss it.
Emergence and the refusal of a fixed endpoint. ALife systems are frequently designed to produce emergent behavior — outcomes that arise from the interaction of simple rules but were not directly specified by them, and that can surprise their own designers (Whitelaw, 2004). This is structurally close to Lee's account of interpretation as the pursuit of "what if?" rather than "what is" (Lee, 2024, p. 22): both value processes whose value lies in not being fully pre-specified by their starting conditions. If there is a musical analogue worth considering, it could be not only performances (our outcomes, sometimes surprising), but also the training (where we learn our interaction rules) that leads to the skill of performing.
Embodied skill on par with real intelligence. Penny argues that skill and intelligence are not two separate things — one bodily, one mental — but a single, non-dual capacity that western culture has arbitrarily split (Penny, 2020). He traces this split to a Cartesian inheritance embedded in digital culture's own founding assumptions, and warns of a broader trend of "de-skilling and dumbing-down" as touchscreens and pre-processed interfaces strip out the "sensorimotor precision and physical effort" that skilled practice has always required. His examples run from flint-knapping to blacksmithing to violin playing. (Penny, 2020, pp. 4–5). If we accept there is "no principled separation between skill… and intelligence" (Penny, 2020, p. 4), then a musician's years of trained embodied judgment, i.e., the accumulated, non-verbal knowledge of a bow arm, an embouchure, a drum stroke's exact application of force, are not a lesser, mechanical poor-relation of the "real," cognitive work of interpretation (or composition, for that matter). They are part of the same capacity, and disembodying that capacity into pure information processing is not simply hard for a machine to do, it may be a category error to expect it. This gives the earlier appeal to "liveness" a more specific description: what a listener may be responding to is not presence or relationship for its own sake, but the perceptible trace of an embodied, effortful, risk-bearing skill being exercised in real time.
Viewed together, these ideas suggest that what a musician can most defensibly rely on is not a fixed property (human-ness, presence) but an ongoing practice: cultivating and displaying real-time, embodied, consequential risk-taking that keeps genuine emergence (Lee's "what if?") alive within a live relationship with an audience, in Small's sense.
However, emergent, contingent behavior is not a uniquely biological achievement; it is a design goal actively pursued throughout ALife and generative AI research, and reinforcement-learning and generative systems already produce outputs their own designers cannot fully predict. "Embodiment," similarly, is not exclusively human; robotics and embodied-AI research treat sensorimotor grounding as an engineering target. Nothing here guarantees that emergent, embodied-seeming machine performance stays permanently out of reach — only that it is not yet routine, cheap, or convincing at the level Trost et al.'s or Ansani et al.'s human performers currently achieve. A realistic view is that musicians currently hold the lead on a moving target.
6. What Would a Musical Education Built for This Look Like?
If the moving target scenario described above is roughly correct, music education oriented toward irreplaceability, rather than toward competent reproduction, would need to look different from the conservatory model in a few ways, without abandoning technical training. Penny's argument, after all, is that embodied skill and interpretive intelligence are the same capacity, not that skill training should be cut to make room for something more "cerebral."
First, it would treat interpretive risk-taking as a skill to be deliberately trained and assessed, not as an ineffable gift some students happen to have. This could be something closer to what Lee's reading of Gould implies: a disciplined, historically informed judgment about when a convention has calcified and is worth resisting, taught alongside, not instead of, the technique required to resist it convincingly. Second, it would take seriously Small's point that a performance's meaning lives in the whole set of relationships in the room, which argues for training that treats audience relationship and real-time responsiveness (the "spectactor" dynamic Penny describes) as a core performing skill, not an extracurricular add-on to "real" musicianship. Third, it would still make room for a baseline of technological literacy, not to turn every musician into a sound engineer, but so that performers acquire the vocabulary to express what they want from, and how they can contribute to, the systems being built around them rather than passively accepting a systems' assumptions.
This is, perhaps, at best, a bet about which currently-human capacities are hardest, not impossible, to automate, and a bet about where a society that wants a live, lively musical culture, should be putting its instructional (and institutional) resources.
7. An Open Ending
So far I haven't directly answered the question "what are we"; that is, what is it that performing musicians posses that cannot be (easily) replaced. But I have explored potential musical possibilities. Exactitude was disqualified immediately, sequencers, MIDI, and even old-fashioned piano rolls have seen to that. Human-ness turned out to rest more on belief whether we are witnessing human performance, live or recorded. Liveness can offer promise if we view it as relational, but still not something an average working musician can rely on economically. Interpretation-as-provocation is compelling but not perhaps off-limits to a sufficiently ambitious machine (or its designer). ALife's vocabulary of emergence and embodiment suggest some concepts that musicians could artistically exploit. For now, the edge performing musicians have is the perceptible, risk-bearing trace of an embodied skill being exercised, inimitably, in relationship with people in a shared performance space. This may seem self-explanatory for those of us who are lucky enough to still do this for a living, but I think it is worth explicitly stating as our state-of-art now. How long we keep this edge relies on many factors: the designs of AI developers, audiences, funding organizations, and ourselves, as long as we continue to search and develop the irreplaceable "what" that we have to offer.
References
Ansani, A., Koehler, F., Giombini, L., Hämäläinen, M., Meng, C., Marini, M., & Saarikallio, S. (2025). AI performer bias: Listeners like music less when they think it was performed by an AI. Empirical Studies of the Arts, 43(2), 1137–1161. https://doi.org/10.1177/02762374241308807
Artaud, A. (1958). The theatre and its double (M. C. Richards, Trans.). Grove Press. (Original work published 1938)
Dorin, A. (2015). Artificial life art, creativity, and techno-hybridization. Artificial Life, 21(3), 261–270. https://doi.org/10.1162/ARTL_e_00166
Dorin, A., & Stepney, S. (2024). What is artificial life today, and where should it go? Artificial Life, 30(1), 1–15. https://doi.org/10.1162/artl_e_00435
Langton, Christopher G. (1989). “Preface.” In Christopher G. Langton (ed.), Artificial Life: Proceedings of an Interdisciplinary Workshop on the Synthesis and Simulation of Living Systems. Santa Fe Institute Studies in the Sciences of Complexity, Vol. VI. Reading, MA: Addison-Wesley, pp. xv–xxvi. [check this!]
Lee, A. (2024). Playing: Or, the subversive politics of interpretation in music. Vantage, 9(3), 22–23; continued in Vantage, 10(1), 26–27.
Levin, M. (2025). Artificial intelligences: A bridge toward diverse intelligence and humanity's future. Advanced Intelligent Systems, 7, Article 2401034. https://doi.org/10.1002/aisy.202401034
Maturana, H. R., Cohen, R. S., & Wartofsky, M. W. (1980). Autopoiesis and Cognition: The Realization of the Living. Springer Netherlands.
Penny, S. (2020). Embodied cognition, digital cultures and sensorimotor debility. Proceedings of ISEA2020.
Penny, S. (in press). Designing behavior: Interaction, cognition, biology and AI. In Encyclopedia of new media art. Bloomsbury. (Preview: https://simonpenny.net/2020Writings/DesigningBehavior.pdf)
Small, C. (1998). Musicking: The meanings of performing and listening. Wesleyan University Press.
Trost, W., Trevor, C., Fernandez, N., Steiner, F., & Frühholz, S. (2024). Live music stimulates the affective brain and emotionally entrains listeners in real time. Proceedings of the National Academy of Sciences, 121(10), Article e2316306121. https://doi.org/10.1073/pnas.2316306121
Whitelaw, M. (2004). Metacreation: Art and artificial life. MIT Press.
1The study of systems that includes the observer as part of the system being observed.
2https://www.stms-lab.fr/projects/pages/somax2/#header
3Lewis, George E. “Too Many Notes: Computers, Complexity and Culture inVoyager.”Leonardo Music Journal10 (December 2000): 33–39.https://doi.org/10.1162/096112100570585.