Toe-to-Toe with Ingo
Toe-to-Toe with Ingo is a podcast from the National Center for Voice and Speech featuring conversations between renowned voice scientist Dr. Ingo Titze and leading researchers and clinicians in voice and speech science. Each episode explores the fundamental science behind voice production, laryngology, and speech through thoughtful, candid discussion. The goal is to advance understanding in the field through rigorous inquiry, respectful disagreement, and the exchange of new ideas.
Toe-to-Toe with Ingo
Perception vs. Acoustics in Voice Training
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
In this episode of Toe-to-Toe with Ingo, voice scientist Dr. Ingo Titze is joined by renowned voice pedagogue Kenneth Bozeman and voice researcher Dr. Ian Howell for an in-depth discussion on the relationship between perception and acoustics in voice training. Together, they examine how singers, teachers, and researchers interpret the sounds they hear and why our perceptual experience doesn’t always align with the underlying acoustic and physiological realities of voice production.
The conversation explores a wide range of topics at the intersection of voice science and vocal pedagogy, including pitch perception, harmonic structure, vocal fry, perceived loudness, sound pressure level, acoustic “buzziness,” and the strengths and limitations of the source-filter model. Along the way, the panel considers how scientific measurement and subjective experience each contribute to effective voice teaching and how misconceptions can arise when the two are confused.
Throughout the discussion, the guests tackle fundamental questions that continue to shape both research and pedagogy: Why do we sometimes hear a pitch that isn’t physically present? Why is vocal fry often better understood as a perceptual phenomenon than a mechanical one? And how should singers and teachers reconcile what they hear, what they feel, and what acoustic measurements reveal?
As with every episode of Toe-to-Toe with Ingo, the conversation embraces rigorous scientific inquiry, respectful disagreement, and the open exchange of ideas. Whether you’re a voice scientist, clinician, teacher, student, or professional singer, this discussion offers a thoughtful exploration of how perception and acoustics together shape our understanding of the human voice.
Over on the road. Must have a code. Last year lift and so become yourself.
SPEAKER_03Welcome to Toe-to-Tee Withing Guides, a podcast from the National Center for Voice and Speech. Where evidence matters, questions are welcome, and disagreement drives discovery. This series brings together leading scientists in voice and speech science for candid, thoughtful discussions around foundational questions in the field. These conversations are intentionally rigorous, occasionally uncomfortable, and always grounded in respect for evidence, curiosity, and one another. I'm Andrew Parker, Communications Director at NCVS, and I'll be serving as your moderator today. Toe-to-toe with Ingo is designed to be conversational. The goal is not consensus for its own sake, but progress through questioning, disagreement, and careful reasoning.
unknownDr.
SPEAKER_03Titsa has been clear that this podcast should not shy away from differing perspectives. In fact, we believe that respectful disagreement is often one of the most productive forces in scientific advancement. Today's episode focuses on perception versus acoustics, a topic that sits at the intersection of voice science, vocal pedagogy, and the complex relationship between what we hear and what the voice is actually doing. Our host is Dr. Ingo Titsa, distinguished professor emeritus at the University of Utah, founder of the National Center for Voice and Speech, and a leading figure in modern voice science. Joining him today are two distinguished guests, Kenneth Bozeman and Ian Howell. Gentlemen, welcome. Thank you. We're glad to have you here, and I'm going to turn the time over to our host, Dr. Titsa, to introduce you.
SPEAKER_01Hello, everyone. It's a pleasure to be together and to have a wonderful topic that we can discuss today. My guests are Ken Bozeman. He's an emeritus professor of music at Lawrence University Conservatory in Wisconsin. And he received his degrees from Baylor University and University of Arizona. And Ian Howell received his DMA from the New England Conservatory, where he also then taught for at least a decade on the faculty. And he is now the, I guess, the chairman of an organization that he himself founded. It's called Embodied Music Lab. And it's a privilege and lab private organization, right?
SPEAKER_04Yes, I'm in my entrepreneurship phase.
SPEAKER_01Okay. Both of these gentlemen have uh tremendous respect across the world, have written books and papers in both uh pedagogy as well as the science area. Um, and uh it's a great pleasure to have the opportunity to just chat together and see what we can uh learn from each other. Um again, our moderator will be Andrew Parker. He's the communication director of the National Center for Voice and Speech. Um and we'll turn the time back to him. He's going to uh shout out the questions and we will jump in as we see fit. Thank you.
SPEAKER_03Sounds great. Well, let's go ahead and uh begin our discussion uh with this thought here. So many of the concepts that singers and teachers work with every day are rooted in both measurable acoustic events and of course subjective perceptual experiences. And sometimes those things align beautifully, and of course, sometimes they don't. So let's start there with our first question. Why do we sometimes hear a fundamental frequency when only higher harmonics are present? Uh Ian, why don't we start with you on that?
SPEAKER_04Yeah, sure. So, I mean, this is a psychoacoustical phenomenon called the missing fundamental, which has been documented for more than a century at this point. And experientially you find it in music treatises, going back even further. And certainly any pipe organ builder understands that the way you get the lowest fundamental is with two shorter pipes that are a fifth apart, right? So conceptually, this has been around for a while, and there's a number of ways that uh people will explain it. Uh, one of the things that I have on my mind pretty constantly at this point is um is trying to be able to like see the conceptual lenses we're using to understand acoustics and to understand perception. One of those conceptual lenses is to understand periodic sound through the mathematical output of a discrete or fast Fourier transform. And so if one is taking that lens toward this question, then the answer then becomes well, the way that the hearing process works and the hearing brain responds to a sound that is full of higher harmonic content, but perhaps missing those slowest components, the fundamental and and you know, even higher harmonics than the fundamental, is that the brain extracts the mathematical difference in hertz in frequency between those higher harmonics. Periodic sound being what it is, we would represent it with equally spaced harmonics in a frequency space. Um, I I think it's really interesting to look at that same question through a different lens, which is essentially, you know, that that Fourier transform is looking at a time domain phenomena that's you know, it's pressure changes over time, essentially. And um it's really interesting. I think in voice pedagogy, especially voice science, much less so, but in voice pedagogy, we don't spend a ton of time looking at waveforms. Um, Ingo, you you you introduce voice acoustics through the waveform in principles of voice production, which I think is amazing. And um, and then you rightly point out that waveforms are highly redundant, and so it makes sense to move into a sort of a summary mode to understand it. And that is absolutely true. But if one removes that fundamental frequency, you know, do a Fourier transform, zero it out in the spectral data, and do an inverse filtered uh transform back to a waveform, the the periodicity is still there. Like I think that's the interesting story with this question, is that you have to remove a lot of harmonics before you get to the point where the sound itself is either so chaotic or the sound itself is low enough in amplitude that we stop hearing that experience of the pitch. One of the things that I think is fascinating, and this falls apart at the edges. So if you really drive a cochlea hard, this is not true. But the pitch that we hear tends to be related to the frequency, sorry, not the pitch, the the tone color of the pitch that we hear tends to be related to the frequency range of the harmonics that we are presented with. And um, and so I think that that is the thing that tells me it's a phantom experience. Because if we were actually being presented with energy oscillating at that slower fundamental frequency, we would hear like a tonal color world that aligned with that frequency. So that's how I think about it. I and there's more than one way to even teach people this basic concept. Certainly just play them sounds. Like that's the way to get the idea across, and it freaks people out every single time you do it. Um, but I I personally, even though most people present it through a frequency domain lens, I think the fascinating one is a time domain lens.
SPEAKER_01There's a a term that hasn't made it into our community very much. It's called the greatest common divisor between the frequencies. So let's give an example. If you had two frequencies, one was uh 400 hertz and the other one was a thousand hertz, and those are the only two frequencies you had, the common divisor would be 200 hertz. Okay, you can divide a thousand by two hundred and get uh an integer number, and you can take four hundred and divide by uh two hundred and get an integer number. So that means that the common divisor between 400 and 1,000 is 200. And that is what we hear. And uh and that is also what the pitch extraction algorithms that are available now everywhere uh try to do. They try to find the greatest common divisor, not just what you see in the spectrum, but going into that concept of uh the brain actually does that division, as you have said earlier, and produces that fundamental frequency. And this is why we still hear the fundamental frequency over the telephone, uh, where often frequencies below 300 Hertz are not even present.
SPEAKER_04They're they're filtered out, but we still hear the fundamental or a baritone singing in front of an orchestra where probably very little of that lowest frequency information is even making it past the proscenium.
SPEAKER_02Yep. Very brief uh addition. I I I think of it like Inga was just describing, that the common denominator, which would be the greatest. Uh and if you if you took a bunch of sine waves, if you're thinking in that version, that are there with the missing fundamental, and you do them additively, they will create a pressure wave that repeats at the rate of that missing fundamental frequency because it's the lowest or the greatest common denominator. So we perceive the periodicity of that sound wave still stays at the fundamental frequency rate, which we perceive then as the pitch. So same, same thing. Um and and as as been pointed out, our hearing is actually very weak lower. We you know, we perceive this actually comes up a little bit a later question. Our our we perceive vibration with two different senses, touch and hearing. And hearing is most sensitive from a musician's point of view in the top octave of the piano, you know, two to four thousand hertz up in that area, and feeling is more sensitive around middle C, around 250 Hertz. So um we if you played, if you played that that fundamental frequency for most pitches that that many of us sing, it's barely audible, barely audible to our sense of hearing. So we're imposing the higher frequency content to that fundamental frequency in terms of timbre and other things in anyway.
SPEAKER_01That's yeah, with another sense, yeah. Nice description, yeah.
SPEAKER_03Great. Well, let's let's move to the next question. Uh uh Ingo, I'm gonna start with you if that's okay. Um the next question we've got is why do scientists use the term pitch extraction when determining fundamental frequency?
SPEAKER_01I was actually angry about that for many years, because we all know that pitch belongs to perception and fundamental frequency belongs to measurement, acoustic measurement. But why do so many scientists maintain the word pitch? It's not only because it's a uh one-syllable word and easy to write, but it's also because they're trying to actually uh get at what we were just talking about, the fundamental frequency that we perceive as individuals, even though it may not be there. And so we have looked at all the algorithms uh that are currently available, whether it's uh uh Pratt or Vocevista or many others, and they all end uh do that uh uh measurement of the greatest common divisor that we've just been talking about. Um and so the the scientific community has gotten uh in tune with us, so to speak, to get the best they can with instrumentation from what we actually hear. And that's why I'm not so upset now that they use the word pitch, because they're trying to get to what we perceive.
SPEAKER_04I I think there's some there's some edge cases that that make it tricky and really kind of poke at at what you're saying that pitch is a perceptual phenomenon. Because if you if you listen to a, you know, as as close to a sine wave as can exist in the natural world, like if you listen to a really simple oscillation um and change its intensity, the intensity of the output, our perception of what the pitch is is gonna change. And it'll bend up in one direction and bend down in the other direction. But as a physical thing, the way that a pitch tracking algorithm would measure it would measure no change. And so, like there are like interesting edge cases like that that, or think about a phenomenon like subharmonics, when those pop into a speaking voice, like there are liminal thresholds where you'll hear the higher pitch and then you'll hear the lower pitch, and like at what point do you move from one to the other? Is like I'm not even sure how one would answer that question.
SPEAKER_01Well, we we just published a study on that, and uh we found that at least informally we hear it uh when uh the subharmonic is greater than 30 percent of the fundamental of the harmonic series. Well, uh if I knew that's when the lower pitch starts to come in and starts to dominate.
SPEAKER_04I would have asked for more if I knew I was gonna get my questions answered. That's great.
SPEAKER_03Some of your thoughts. Oh, no, sorry.
SPEAKER_04I was I was just gonna say, Ingo, I um I I have a decent understanding of pitch tracking algorithms, and I I work really actively in Pratt, and I know, for example, the difference between like a raw cross-correlation approach versus a filtered autocorrelation approach. And and I understand that the the autocorrelation approach transforms things into a Fourier transform domain, right? And then it I think looks for the alignments of the harmonics like we're talking about, and then it inverse filters it, but then a raw cross-correlation is really sliding the waveform against itself, looking for alignment. Is that is that right?
SPEAKER_01Yeah, I think that's that's exactly what's going on. Yeah.
SPEAKER_03Cool. Thank you. Um, Ken, what are some of your thoughts?
SPEAKER_02I'm I'm good with what we said already. A lot of this is above my pay grade, so let's move ahead.
SPEAKER_03Fair enough. Fair enough. Um well uh let's uh uh um uh let's come back to uh to this next question here. Um and Ken, maybe I'll start with you here. Why are low frequency vocalizations perceived as rougher than high frequency vocal vocalizations?
SPEAKER_02So if you're singing, if I use the word pitch, if you're singing a low pitch, a low sung pitch, and you have good chord closure, good rapid chord closure, you have a lot of high frequency spectral content in the signal that's going on. And high frequency spectral content, if it has higher numbered harmonics that are close enough together, they start beating against each other to our perception and introduce a buzziness, an auditory roughness. But it's it's a periodic roughness, not the roughness of asymmetry or pathology, which is a noise kind of roughness. So it's a it's a it's a we in the pedagogy world, we consider it a good roughness. When you sing in a high pitch, the spacing of the potential lower harmonics of that pitch are wide enough uh relative to our hearing sensitivity that they don't beat against each other. And those high, which you need to have auditory roughness from harmonic content, it has to be high, closely clustered and resonated harmonics that trend, you know, put in the signal these closely clustered beating against each other harmonics. When you sing a really high pitch, those that level of the uh spectrum is way beyond our hearing sensitivity, well off the keyboard, and it may also not be being resonated by any higher resonances of the vocal tract. So you just hear the more or less smoother lower end of the spectrum where the harmonics are well spaced, and you just get this pure sort of round, warm color. Thank you, Ian Howell, for that information. Whereas with the uh lower, so bases are really buzzy because they're singing low pitches with a lot of resonated, high clustered harmonic content. Sopranos are really smooth, and every voice type migrates from relatively buzzier low pitches to relatively smoother high pitches, and basses get from really buzzy to sort of buzzy. Sopranos go from sort of buzzy to really smooth, and everybody else is somewhat in between in terms of where their range lands with these things. So perceptually, that's that's how that works. And you can use that information pedagogically to actually indirectly train laryngeal function. If the output target goes from buzzier to smoother, that tends to invite the better function from the larynx. We get appropriate levels of closure. Uh the other piece to it is that shorter, thicker cords that close with more contact and for longer each cycle generate more high frequency content. Longer, thinner cords that have a shorter uh close quotient uh tended to generate a steeper spectral slope with fewer high spectral contact or less high spectral content, which also contributes to this buzzier to smoother uh across range uh percept that we that we have.
SPEAKER_03Uh Ingo, uh what are some thoughts you had on that?
SPEAKER_01Uh just to add to that, that's very, very good. Um, Ken, you you summarized it very well. Uh in the whole uh area of sound perception by humans, there's something known as a critical band. A critical band is a band of frequencies in which the cochlear uh mechanics can no longer isolate one pitch from another. Um and so if two harmonics or two any two frequency end up in one critical band, then what we perceive basically is how these two frequencies fight each other as opposed to uh hearing them individually. And uh critical bands can be well uh defined uh over a whole range of frequencies. Um I love uh the book that we no longer use, but uh Harvey Fletcher uh was uh a pioneer in our field. Uh he actually was the first one to establish an acoustics laboratory at uh Bell Telephone Laboratories after World War II. And uh his book is called Speech and Hearing in Communication, and he always puts the two together. He always has what we perceive with our auditory system together with what we produce with our vocal system. And um, anyway, critical band theory is uh important factor in in all of this discussion.
SPEAKER_04I I've read that book cover to cover several times, and and beyond the information being not only interesting and good, and also being part of a really fascinating historical sweep, he is a hell of a writer.
SPEAKER_01Yeah.
SPEAKER_04My god, what like what lucid prose about complex ideas? Just beautiful.
SPEAKER_01Yeah.
SPEAKER_04True. Um, I as a as a performing artist, just to throw it out there too, that there's always the interesting question about like what why would we perceive auditory roughness in a voice at this pitch versus this pitch? Um and you know, we tend to use this shorthand, like Sundberg uses this in Science of the Singing Voice, where we say any two harmonics that would fall within, you know, basically a little bit narrower than an equal-tempered major third will will trigger auditory roughness. And that's that's a simplification, that's a good simplification. And we could understand that as the harmonics go lower in frequency, that critical band actually gets larger, and as they go higher, it gets narrower. So, like there, anybody who wants to jump off the plank into that world, there's more interesting stuff to look at. But singing with other people, like that opens up some real options because all of a sudden you can load your your own cochlea with frequency information that you couldn't make yourself. And anybody who has sung close harmony understands that like there are these intervals that you will hit where it is as though the sound is a solid thing in the room. And it's be I think it's because you know the the human brain's response to the cochlea being stimulated within critical bands is like it's it's that's how we experience it. It would be like sweet versus sour, right? It's just like a a sense a sensory experience of that physical phenomenon. Um, but any anybody who has sung close harmony go sing a whole step with another human being in a stairwell. And like I feel like that gets the idea across almost better than anything else.
SPEAKER_03That's great. That's great. I'll have to remember to do that. Um all right, let's let's go to our next question here. Um, uh because we're talking a little bit about buzziness. So I think this next question really kind of uh uh helps with this. Um And um Ken, I might come back to you to start here. How much of our pedagogical language conflates the sensory experiences that we informally describe as fuzziness?
SPEAKER_02Conflation in I think involves conflating two different things at least. You can't conflate something with itself. So I think we're conflating our hearing sense of frequency with our feeling sense. So our vibratory senses peak around middle sea. And our our colleague that we co-teach with Chadley Ballantine has really brought this information to the fore from hearing impaired research. How do people with no hearing or very low hearing experience sound? They experience it as vibration and it peaks around 250 Hertz and then it tapers off. And above a thousand hertz, we cannot feel vibration. We can only feel vibration up to a thousand hertz. That's roughly soprano high C. We're just beginning to get into the strength of our hearing perception at a thousand hertz. It really peaks somewhere probably near around 3,000 hertz in the top of the piano. So the difference between those two perceptually causes us to conflate those experiences. So for example, singers generally, particularly non-treble singers, basses, baritones, and tenors, and even the lower octave of treble voices, need high frequency content in their spectrum to carry over orchestras. Why do they need it? Because high spectrum content in the top occur of the piano is where our hearing is sensitive. And if we have that there, it'll cut over the spectrum of the orchestral background and so forth. When that is present strongly in, say a tenor's high note or a bearitone's high note, we feel vibration that we think we're feeling the content of that high frequency content. We are not. We are feeling strengthened low frequency content. But we associate it because it's correlated with that high frequency content. Because if you strengthen any part of the signal, my understanding, please correct me if I'm wrong on this. If you strengthen any part of the signal, it's like the whole signal is stronger. If you strengthen the high part, so we we feel the low frequency content buzzing away when that high frequency content is present, and we conflate and think we're feeling the high frequency content. We're not, but it's still actually useful biofeedback because it's there when we get a strong signal. And so we conflate those things. Um that's at least a start into this equation. Well, we're we feel low frequency content as feelable vibration, we hear high frequency content as a ringing, you know, a high buzzy timbre that Italians call squillo, and we have various terms in the history of pedagogy for it. At least that's a a stab at this subject.
SPEAKER_04I I will make the next stab. Um, I I think I'll say a controversial thing. I think that a lot of voice teaching is really challenging to falsify. And uh and this is one of the things that I think is kind of a home run, is what Ken is talking about. And so, for example, I there's no way to gather evidence about this. This is just my lived experience as a singer and a voice teacher. But any voice teacher I've encountered who has tried to get a singer to experience some sort of forward vibratory sensation, and you know, we can certainly talk about whether that's a good pedagogical approach or not, but they always use bright vowels for that. So they'll use an an A or an E kinds of sounds, like snarly sounding vowels. And I think people have this association that the feeling is caused by the brightness that is present in those vowels that is not present in like an ooh, for example. And um, and it's it's exactly backwards. It's actually the the monumental magnitude of the first formants of those vowels that is almost certainly generating the sensory experience that is like a vibrotactile experience. Um, bridging these things is bone conduction hearing, like that it's a it's a fascinating bridge between uh airborne hearing and our sense of vibration. And part of me does wonder how much conflation there is in people's minds that they they think they're feeling, but actually they are hearing. Um, Chadley and I actually just finished up a study which looks at this. We were basically doing like audio audiograms. Like if you were to go to an audiologist and they would check your hearing, except we did it with um playing people's own voices back to them in various parts of their body with contact like vibration speakers. And um and it is really interesting. Like the threat, the threshold for feeling vibration as a sensation, like Ken said, is shockingly low in the frequency domain, given what I think most singers think is causing that feeling of vibration.
SPEAKER_01And even lower than that is the actual resonance frequency of the tissues that are vibrating, uh not just the receptors, which you say are tuned to around 250, but the resonance frequency of buckle tissue and other tissue uh in our um system is more like 120 Hertz or so. Um and so when we do uh resonant voice therapy or anything with semi-occluded vocal tract, we always say it has to be easy and has to be buzzy. But it's good for speech because that's where we are. We're in 100 and 200 hertz. By the time you get up to 500 hertz, or as you mentioned, a thousand hertz, none of those sensations are there anymore.
SPEAKER_04You've got to be really loud to have those sensations.
SPEAKER_03That's great. I appreciate the discussion here. Uh let's uh let's go to our next question. And uh Ingo, I'd like to start with you. Um, why is vocal fry more of a perceptual phenomenon than a mechanistic one?
SPEAKER_01Well, uh from what I have learned, again, from perception of low frequencies, is that when frequencies get somewhere 70 or 50 hertz or lower, then um those frequencies aren't perceived by our brain anymore as places on the cochlea, but rather as individual events that are sent uh to the auditory nerve and finally to the brain. In other words, what you hear is just pup, pup, pup, pup sounds, individual sounds that aren't necessarily connected in a continuum the way normal sound is. And and that the the breakover uh region there is around 70 hertz or so. And but that means that you can have many, many mechanisms that produce the same thing, as long as there is the perception of sound, gap, sound, gap, you can use all kinds of different um mechanisms inside in the vocal folds and still produce that same effect, that same feeling of gap, sound, no gap, or sound, no sound, the sound, no sound. And that's why I don't think it's a mechanism. And uh the people who put M's in front of mechanisms like M1 and M2 and so forth use M0 for Fry, and I have not yet been convinced that there's any one mechanism that produces this, but there are many, many mechanisms that give you that perception.
SPEAKER_04I think there's um in the in the metal screamo distorted singing practice world, there there's also really interesting use of the term fry. Like it it nests as a compound term, so there's something called a fry scream, for example, that um would definitely be considered to be pitched, that has nothing to do, I think, with with what most of us would call vocal fries. So, like I would say even uptake of the term itself is is fracturing at this point in a really interesting way. Interesting because it points us towards new new sounds maybe we didn't know the human voice could make. Um, I want to ask you a question though, Ingo. I I always wonder about this like how much of that 70 hertz-ish cutoff is because of the periodicity, and how much of it is because the nature of the vocal tract when excited is to dampen rapidly and especially dampen higher frequency energy, probably more rapidly than lower frequency energy. How much of it is the gap? Which is to say, if my vocal tract were like a wine glass and I pulsed at 70 hertz and the resonance just kept going, you know, for several seconds, would we still perceive it in the way that we perceive it now?
SPEAKER_01Well, in terms of the limited experiments that I did, and that was many years ago with uh Anat Kedar, um what we did is we simply changed the length of the gap where there was no sound uh and left everything the same. Uh in other words, the pulse that uh that was created by collision of the vocal fold and then the resonance stopped and then we changed the gap, and that's how we determined 70 hertz was the transition frequency. But you're right, if uh if you figure out a way of not dampening uh the energy in the vocal tract easily, even at 70 hertz, you could still have a continuous sound, and that probably then would not end up being fry-like. Uh so you could have normal uh vibration um in whatever register you want to call it, maybe down to 50 hertz or even lower.
SPEAKER_04Very cool.
SPEAKER_01So I I I think your your question is right on target.
SPEAKER_03And I'd like to get your thoughts on some of this.
SPEAKER_02Very, very few. First of all, ignore if you're right, if there's no mechanism there, then maybe the zero is an appropriate mechanism zero. There is no mechanism.
SPEAKER_01Okay, I get it. Zero means no mechanism. Okay. Right. I got it.
SPEAKER_02Right. So the other one was uh you can take a click, a spaced click. I I I make this comment. I say pressure variations are experienced by humans on a on a continuum. If they're really slow, we experience that as weather. Yeah. If they are if they are abrupt but infrequent enough, we perceive that as rhythm. But if you get them close enough together, then we start perceiving it as pitch. So it may be, I'd be very curious to see, because you we we've got a I've got a little app that does this where that Christopher Besh put together for us, actually, as a colleague, where he starts with clicks and he speeds up the frequency of the clicks, and you hear them as individual rhythmic clicks, and Benji gets fast enough, all of a sudden you start hearing them as pitch. And if you run that through a Fourier transform, you'll start getting frequency, frequencies appearing, but you don't get the frequencies from the individual clicks, even though every sound has frequency content. You just get this vertical striation, you know, and then anyway, it's very, very fun. So then, and then if you get them fast enough, fast enough, then they're supersonic and we can't hear them at all. So we go from weather to rhythm to pitch to supersonic.
SPEAKER_01All right. Okay.
SPEAKER_03I uh I'm I'm I'm really enjoying the discussion. So I uh we've got a couple more questions here that I'd like to get to here. So um let's uh maybe move away from some of the mechanistic discussion and talk about sound pressure. Um and Ingo, I'll I'll start with you again here. Why does sound pressure level not always correspond to perceived loudness?
SPEAKER_01Well, I'd bring us right back to Harvey Fletcher. There is something known as the Fletcher-Munson curves, and they uh equal loudness curves. Um everybody in our field should know that diagram. Uh that should be the number one uh along with maybe uh uh formant uh frequencies or whatever of vowels. But it it shows that uh our ears are most sensitive to frequencies um around a thousand to three thousand Hertz. And um if you want to get the same loudness perception uh at lower frequencies and higher frequencies, then you have to give more sound pressure level to get the same uh effective loudness. Um so um but this is not how a sound level meter is calibrated. We it can be calibrated to have that same uh difference uh that we have from our auditory system. But most of the time um the sound level meters just gives us the amount of sound physically that is there. But our auditory system gives us basically what is known as the Fletcher-Munson curves. Um yeah, maybe Ian you can add to that or everybody else.
SPEAKER_04I think it's a little bit like asking a question as to what color a flower is, and then it depends on whether you're a bird or a human, because you know, we we can't see ultraviolet, for example, and some colors radiate in beautiful ultraviolet color. And so the I think the entire world that we know is mediated through our senses, like like we everything is beyond a veil that we can't really pierce, and um, and so you know, the the Fletcher Munson Equal Loudness Contours is like a great example of this, um, because we shouldn't anticipate being able to hear certain frequencies. Well, this is so true, it's in the law, right? Like OSHA and NIOSH requirements for work sites, they can have much higher intensity sounds as long as they fall outside of the range of hearing sensitivity. Um, and there's there's a low frequency world going on outside right now that is cacophonous that we're just not we're not aware of and we'll never be aware of. Um, I I think I think this is another thing that voice teachers really need to know about, and not even from a mathematical model point of view, but just from a practical application point of view. Because if if you're a voice teacher and you teach in an eight by eight studio, you are significantly closer to your singer than you will be if they're an acoustic singer than you will be when they are on a stage. And proximity to a sound source, because the inverse square law, right? Proximity to a sound source, it's going to be more and more intense the closer you are. And one of the interesting things about those contours is that the more intense the sound is, the flatter they are. And so the implication of this is if you are trying to train someone to just like sing the bejees out of music on a stage and put out a lot of perceptible sound from a distance, and you are not overwhelmed by the warmth and brightness in the studio that will then be attenuated in the concert hall, then you're actually not training them to make the sounds that they need to make. And and so this is this is a strong argument, I think, and this is not a podcast about voice teachers, but but like it's a strong argument for teaching some of your lessons in a recital hall and sitting 40 feet away from people because those curves kick in and it makes a big difference than the sound a singer makes.
SPEAKER_01And the and directivity of the sound is something very important, uh, particularly in the animal world. We just studied uh the vocalization of the P-haw bird, which is a uh small bird uh in South America, um lives high in the canopy uh area, and uh that bird um they have recorded of producing 110 decibels, presumably at one meter distance. I say presumably, but that's impossible. I went through and I computed the caloric intake of that bird, and and the it would the amount of wattage that you get out of the mouth is more than what that bird takes in during the thing. And so what's happened is people couldn't get close enough to the bird to really put a sound level meter right near its uh beak, but they uh uh uh uh extract it backwards, thinking that it was perfectly uh, you know, uh spherical sound. But it isn't. It was very directed, and there, therefore, they got that high number. But that that that also happens in in the in the hall. I mean, uh men say that why do the women always outdo me uh, you know, with the high frequencies uh back deep into the into the uh concert hall is because the high frequencies are more direct, they can be beamed, the low frequencies spread locally and you know don't reach very far.
SPEAKER_04That's why if you're if you're ever staged to sing upstage, you better be like seven inches from the back wall and it better be made of plaster so that your sound will actually bounce out into the hole.
SPEAKER_02And go ahead. A couple of comments. In addition to the threshold of hearing in that curve, which is a wonderful curve, the top of the curve is the threshold of pain. Yeah. So just right where the where the threshold of hearing is the most sensitive, in that sort of, I I say two to four thousand. I see that little, there's a little dip right in the pretty much the top octave of the piano, whether it's one to three or two to four. There's a dip in where it starts hurting if it's that intense, uh, which is actually how toddlers get their parents' attention. They really have a lot of high frequency content right in the area, and that has a lot to do with the dimensions of the uh ear canal, which is also a resonator, which boosts that frequency region from the outside world before it hits the eardrum and even raises it higher than it was in the air before it hits the eardrum. So, yeah, there's that as well. And as you mentioned, the high frequency content is much more linear, which is why the tenors and the the you know non-treble voices need a really strong high frequency component to zing out over everything in the hole. And I have heard reports of people saying as the tenor fans the audience like this and tries to hose them down with the sound, you can hear the loudness, the loudness per se swing by you with almost like a Doppler effect. Yeah.
SPEAKER_01Yep.
SPEAKER_03That's great. That's great. That's right. All right. Uh, a couple more questions here. Um uh Ian, I'll start with you on this one. Yeah. Um, what is acoustically weak may be perceptually strong, and what is acoustically strong may be perceptually weak. So, how does that affect voice training?
SPEAKER_04Yeah, this is one of my favorite things. Um, we've brought it up a little bit when we were discussing uh uh vibrotactile sensation versus hearing sensation. Um, but you know, the the the more that I've learned about acoustics and dug into the scientific literature around it and you know considered what what is the transfer between the flow world and the pressure world, and what rates of change are actually generated in in the vocal tract and above the glottis as we phonate. Um, one, it becomes transparently clear that every acoustic recording that anybody has ever listened to has a high pass filter, right? So we should say that out of the gate. And so we immediately lose pressure changes that are slower than that. But two, what becomes really clear is that the the slowest oscillating resonance of the vocal tract, what we'd call the first formant in a radiated sound, probably, as a physical thing is just a it's a high amplitude oscillation. And to the point that when we sing in a way where the second formant dominates and has a higher amplitude oscillation than the first formant does, like those are those are the edge cases, those are money notes, those are high belt sounds, those are you know Pavarotti or Corelli going up to their high C and doing a specific tuning thing so that they really have a standing wave that's generated by their second format. And so, um, as a physical thing, like I I try to understand how our equal loudness contours distorts our sense of the physical nature of the pressure pattern that the human voice creates. And so, like a E. Or even like a oh, I don't know if this will come through my zoom filter, but like any sort of overtone singing type thing where we hear a really piercing high frequency overtone, it is really present in our percept of the sound. But if you record that and then break out the different components and don't look in a spectrum, which has probably a bias of some sort that is boosting higher frequency components, but you just look in a waveform, the amplitude of that thing is physically small. And it's fascinating to me because if we're looking to leverage nonlinearities and we're looking for interactive qualities between what the vocal folds do in voice production and the response and the behavior of the vocal tract, like the stronger those interactions, the stronger the interaction is. And so I think a lot of the sounds that we think of as being strong because our hearing is so sensitive, are actually weak and probably don't have as much to do with sort of the ongoing contributions to self-sustaining oscillation as we think. And because our ears are insensitive to the slower components, we don't like we have to train ourselves to notice them to an extent. But I'm as a teacher, I'm on a kick at this point to just get everybody to notice the warm part of their sound, because the warm part of their sound is the physically strong part of the sound. And and if we want to build acoustic energy and we're looking to make power, like that's where joy is found, is in that part of the spectrum.
SPEAKER_02I can chime in on that. The warm part of the sound, the lower part of the spectrum uh plays typically, in some sense, the first format, right? It is the first format content. And vowels are made of two different colors of sound as since uh individual frequencies have a tone color. You have a lower, warmer color and lower frequency sound, ooh and o like literally if you isolate them and play them, and higher frequencies have other vowel-like colors, high enough it just sounds like E. So you have this range of other colors. We most of our vowel targets in language come from the higher spectrum content from the second format contribution.
SPEAKER_04Yeah.
SPEAKER_02And I do a I do a childoscope whisper uh exploration that that is helps train resonance tuning. Because if you introduce a whisper, it has uh theoretically every frequency in it. It's like a sweep of frequencies. So it completely populates the transfer function of the vocal tract. And if you and you can hear, then in that case, your resonance structure and the formats you hear on the outside are the same thing because you've completely populated the resonance structure with spectral content and it radiates, and you can hear those. We hear the second form of contribution way stronger because our hearing is stronger at the higher end. So if I do an example, you can possibly hear the second form of content of the that series of vowels. But you can hear the lower one if I do ingressive clicks. And they're they're both there. And you can train your ear to begin to hear them both. Well, hearing that lower one, like Ian was saying, we tend to ignore it for two reasons. It's not the language sound we're going for, it's the complementary sound that completes the vowel and makes it sound nice. But the the the identifying color of the vowel is most for most vowels, except for ooh and o comes from that second format contribution. Not from the I call that I call the higher one the over vowel. I just for the studio I use terms that are friendly. Let's call it the overvowel. I could have called it the higher vowel. And the under the the one, the the other one I call the under vowel. And for the warmth, we need sufficient undervowel in the radiated signal to have a nice balanced kaudoscoro tamer, which is pedagogy term, historic pedagogy term for a balance of high and low frequency content, warmth and ring, warmth and clarity. But but it's a trick to start hearing that that undervowel color.
SPEAKER_01So normally if I do a most people with with good hearing that haven't lost hearing off the top, hear the and they don't hear the so Ken, why don't you just call it the uh higher color and the lower color? Yep, you could.
SPEAKER_02But but it has a bowel-like color that they identify with. But you're right, it's the higher color and the lower color. I could have called it anything. It's too late in. I've already I'm already in print calling it the overval and the undervalued, and now that's spread out, they're spread out through the pedagogy world, but identify it as the second form of contribution and the first form of contribution so they know, you know, more scientifically what's creating it, you know. But in the studio, I don't tell my students, can you give me a little more second harmonica on your first format? I don't do that kind of thing to the poor student in the studio, you know. Can you give me a little bit more of that warmer vowel color, more than higher, you know, the brighter? I actually in the studio I try to talk about the brighter and the darker war uh colors. I give it we need more of the high, brighter and the low darker. So I use a variety of terms for that. But fair enough. I just, you know, I made up another term. I know you didn't care for that when I first presented that.
SPEAKER_03Um all right. Let's uh Ingo, do you have any thoughts on uh on this topic that you'd like to share?
SPEAKER_01Uh no, I think it was well well explained by uh our great, great uh teachers of singing of all styles. So much appreciated.
SPEAKER_03All right. Um I'm gonna come to the next question here. Um uh and uh um uh uh Ingo, we'll we'll we'll start with you. Can we still conceptualize voice production through the linear non-interactive source filter model and then simply tack on ideas related to non-linearities?
SPEAKER_01Generally, no. Amen. We can start there, but we don't end there, particularly now with the uh new uh sound productions that everybody wants to make with multiple sound sources, uh, for example, the two folds and the ventricular folds vibrating at the same time, and maybe even the aerial epiglottic fold, then you have interactions between these sources, uh, and you have the interaction of all the sources with the vocal tract, and that creates actually new frequencies because every time, uh, if you for a moment uh allow me to say a sinusoid is one of the frequencies, every time you multiply a sinusoid by another sinusoid of a different frequency, you get what are known as difference frequencies. There are literally new frequencies that are produced. And the linear theory will never give you that. I don't you can work all your life on it, it will never give you uh that difference frequency. And that's what a lot of the uh heavy metal singers now are counting on, is to have this enormous spectrum. For example, if you have two frequencies, um uh say one of them is 20 hertz and the other one, no, let's say one is 120 hertz and the other one is 100 hertz. The difference between the two is 20 hertz. In a linear system, you would never have that frequency produced, but in a nonlinear system that will actually show up in a spectrogram. It's physically there. And so uh we have to learn how to deal with the interactions and the nonlinear concepts. Uh it's uh sometimes frustrating, but uh it uh it it belongs to our world.
SPEAKER_04I um is it okay if I jump in? Yeah, on this um I'm gonna say things that I think will make people mad. So I'll just preface with that. Um I I come at I come at this question from from like a voice pedagogy classroom teacher point of view. And you know, voice acoustic, one of the things I love about how Ken presents voice acoustics is it is at its core phenomenological. Like it's about sounds and it's about registration things that occur, which can then be explained through acoustical models. Um, but I think a lot of every year there's a group of graduate students in voice who go through voice pedagogy classes and sit through voice acoustics units. And I think come out the other side confused because I think that there are a visual vocabulary that arises out of the linear source filter model, which makes people imagine that the the discretized spectrum from the vocal folds and the spectral domain vocal tract transfer function, and then the radiated combination of those two things, that that is actually describing a characteristic of voice production. That the vocal folds are actually generating a series of independent sign tones, which are then individually acted upon in the vocal tract. And we we see it all over our teaching models. People will say, oh, your pharynx is a Coke bottle and your mouth is a Coke bottle, and they're different pitched coke bottles, and that's why that's why when the pure tones go through those Coke bottles, you know, you get a first formant and a second formant, which I think as a model means well. I don't think that's really grounded in reality in a physical sense necessarily. And um, and I'm not really convinced that the voice pedagogy field has grappled with two quotations, which I would love to share. I'm gonna paraphrase them. One is Gunnar Font, if if you go to his book where he lays out the source filter model as as he amalgamated it, um, one of the things he says is uh paraphrasing, he says it's a quasi-scientific misunderstanding to imagine that the contribution of the vocal folds is a literal set of harmonic overtones. Um, and then he goes on to describe it in this mathematically elegant manner, as though it were, right? And that's the beauty of understanding both sides of that at the same time. Uh, if you go back to Helmholtz's work, his English translator, uh Alexander Ellis, uh, it's it's in the margins, unfortunately, like it's in his editor's notes. But at one point he he writes that um Helmholtz did not actually think that the vocal folds produced harmonic partials, which were individual sign tones. But as his theory of hearing presumed that that's how the ear processed things, it just makes sense to imagine it as though that is what the nature of voice production is. And like the deeper I've gotten into into voice acoustics and really trying to understand nonlinearities as a as a way to understand voice production itself rather than post hoc voice analysis, uh, I have just it has gotten harder and harder to literally imagine that the linear source filter model is accurate. And and so if that is true, what fills that space? Like what model could we use to simply introduce you know voice acoustics to voice pedagogy communities? I kind of think is the model you use in your book, Ingo. I mean, that's that's what I keep coming back to is understanding it as flow pressure in a time domain, and then very having a clear understanding of what it is that a Fourier transform does to that signal so that we don't reify the output of it. Uh, because otherwise we're we're thinking that harmonics are bouncing back when really it's just a resonance that's bouncing back in time or out of time with the oscillation of the vocal folds. So maybe I'll give it my voice scientist card now if that's entirely incorrect. But but I am I am curious about your thoughts.
SPEAKER_02You want me to dive in before Ingo straighten this all out and gets your vocal? Yeah, dive in. Go ahead. So um, well, first of all, I don't think I I taught pedagogy for a number of years, 26 years when I was teaching at Lawrence. And I don't think I ever taught the linear source filter theory, because we've known about nonlinearity for quite a while. So I always, whether or not, you know, and again, some of this is above my pay grade because I I like to tell people uh very honestly, I'm not a scientist. I'm a voice teacher that tries to be science informed. So I will always yield to the scientist to straighten me out if I'm if I'm misrepresenting it. Second thing I want to say is these are models of explanation. And as we know, all the way from the statistician George Box, all models are wrong because you're modeling something that's too complex. The only thing that isn't that isn't that isn't an inaccurate model is the thing itself. And you're trying to explain the thing itself. So you have to do a model which deliberately simplifies and you know eliminates some information. So therefore, we teach models. So I always thought the nonlinear source filter theory as the way of explaining these phenomena, even though phenomena, even though realizing that's also got problems. Um and and uh to pick up on what Ian was saying, the problem with the terminology, one of the problems with the terminology, and I still use this model because I think it's very useful in certain ways of the source filter, is that implies there's an implicit chronology. Source filter. So you have this chronology, and we all know that the only thing the vocal folds do is open and close at the fundamental frequency. Now they can do that opening and closing with some variety, thicker or thinner, longer or shorter uh open or closed phase, uh uh greater or lesser contact quotient. So we have the closed quotient, the contact quotient, and the rapidity of the closure versus the opening. All of those affect the nature of that single fundamental frequency oscillation, which has an impact on the potential spectrum, right? If you have more closure, longer, longer closure per cycle, and quicker closure, you're gonna have more high frequency content in the output. But what is that high frequency content? It isn't the harmonic series in a single pulse. In a single pulse, it's a noise. It's a noise sweep. And if you even in the Fourier transform, if you choose the right bandwidth, you will see an individual pulse is a resonance-shaped noise. It has the incomplete performance structure of the vocal tract because it's basically exciting the air in the vocal tract, which which gets excited relative to its own transfer function, its own, you know, resonance structure. So you'll get a series of pulses that have a resonance structure. What then is being periodically re-it doesn't introduce a set of harmonic frequencies, though they're they're included in the noise frequencies, they're not special at that point. And the only ones that are featured are the ones that land within the bandwidths of those resonances, the critical band that Ingo was talking about earlier, those get particularly strengthened, and the ones in between don't get much much help. Then that noise, that shaped, that resonance-shaped noise is reiterated at the fundamental frequency. And the periodicity of that fundamental frequency then enables only harmonic frequencies to survive in the radiated spectrum, but only those harmonic frequencies that previously landed in each individual pulse within the bandwidth of resonances of the vocal tract. So you have it's interactive from the start, which leads me to say this, which so far I can't get any scientists to agree with me on. So you want some controversy. So here's the controversy. What we normally think of as the source, the vocal fold oscillation, and the filter is the transfer function of the vocal tract, is it turns out in an on the on an individual pulse, impulse of the vocal fold closure, which we all know if we've known since at least the 90s in cook, you could probably straighten me out since probably much earlier, that it's this rapid closure that is the strongest pressure change of the airflow. So that's where the real power is, is in that closure. It excites the vocal tract. I lost my train of thought there. The oh, so the the thing that's being reiterated is that transfer function of the vocal tract. Yeah. And the only harmonics that that get even a chance to be strengthened by the periodicity of the vocal fold oscillation are the ones that landed within that the bandwidths, the the critical bandwidth of the resonances. Then those get filtered further than by the periodicity of the vocal fold opening and closing. So I say, this is the part that's controversial, that the so-called filter, which is the transfer function of the vocal tract, is actually crucial in the source itself. Each individual pulse is the source sound, and each individual pulse is a noise, is a resonance-shaped noise. So then where do we get harmonic content by the periodicity, the ongoing periodicity of that sound only reinforces that subset of harmonics that landed in resonances, the periodicity essentially equals harmonicity. So then the source that we call the source, the periodic oscillation of the vocal folds, is actually filtering those periodic noise-shaped resonances for harmonic frequencies. So the source is the filter and the filter is the source, and they're all so interactive that you can't sort them out. So there's my Cliff Notes version of the R.
SPEAKER_01Based on that, maybe our introductory textbooks in the future should simply start to say that sound is produced in humans and animals with a source airway system. So we never separate the two. We always say that the sound is produced with a source airway system.
SPEAKER_02Yeah, I'm good with that. Absolutely.
SPEAKER_01Yep.
SPEAKER_04I I think one of the things that I appreciate about your working, that I I feel like the vocabulary of the linear source filter doesn't really accommodate, certainly doesn't invite people into, is just like the notion of the flow pulse skewing. Like the interaction of the vocal tract air mass at that time scale seems really important for the sound that the voice ends up making. And and again, if we if we if we separate the vocal folds phenomenologically, like it's actually really difficult to understand that interactive property of the vocal tract.
SPEAKER_01And we should say that you mentioned Gunnar Fant, you know, he and his uh colleague Ananta Padmanaba uh back in the 70s already agreed that this was not a linear combination between two systems, but that there was always feedback between them and they interact with each other. Even though, as you say, in his book, he then goes on and explains everything with a linear concept. But that's because, you know, our brains uh just don't follow these things that quickly. We have to understand it in simple tune uh terms first. And then then we uh add the uh the complexities. Yeah.
SPEAKER_04I think um I think this is a good example of the very first thing that I said, which is that there are frequently multiple lenses to try and encounter a phenomenon through. And um, and so the the condition that I think is really interesting is like how how do we understand a resonance that falls between two harmonics? Because if let's say you have a resonance at 2.5 FO, right? Just imagine that interaction. Um the spectrum, if you have a long wind. Fourier transform the mathematics of that are going to prioritize representing that acoustic energy in terms of multiples of the fundamental frequency, just because the periodicity itself is interacting with every one of those sine cosine like phase shifted pairs to do the testing of the complex signal. Um, and so I think it's really easy to imagine, oh, there's energy at 2 FO and energy at 3 FO. But if you look within a single pulse response, actually there is energy at 2.5 FO. It just keeps getting reinterrupted. It just keeps getting restarted and restarted and restarted. So it never has the ability to build the energy that we would normally see in a standing wave in the vocal tract if a more virtuous alignment took place. And um it's just really fascinating to me, like looking at that as the phenomenon that a waveform represents, it seems fairly obvious there isn't energy at 2FO and 3FO. There is energy at 3.
SPEAKER_01But even the linear theory tells us that the there the resonance has a bandwidth, and therefore it's not just at one frequency, but it's a spreading of an entire uh collection of frequencies in that bandwidth. So if a harmonic lands anywhere in the bandwidth, it's still uh energized.
SPEAKER_04Yeah, I just think it's an interesting um it's an interesting mind puzzle, I guess I would say that. It's an interesting thought experiment.
SPEAKER_01Yeah. Right.
SPEAKER_03Okay. Well, I'd like to start with uh, or I'd like to end maybe with a broader question for all three of you, if and if I can ask if you uh keep your answers brief. Um but uh the uh uh you know for singers, teachers, and researchers, uh what's the single most important distinction between perception and acoustics that you wish more people understood? And and Ken, let's start with you.
SPEAKER_02Um Well, things are rarely as they seem. It's useful to know the difference. And the singer sings with what is called procedural knowledge. Declarative knowledge is the science stuff, the measurable part. Procedural knowledge is what it sounds and feels like to the singer to do it. Realizing that difference and then being guided by a knowledgeable helper, teacher, as to what is what are these interactions going to feel like and seem like to you as the producer of it is actually the knowledge the singer needs to know. Uh I think the teachers should learn the declarative knowledge because it's a conversation between those two that guides their instruction, a very informed conversation between what's actually going on and what does it seem like to you, right? Uh, but but the singer sings with the what it seems like. That's that's the thing. Once they know the feel it feelable path and the hearable path of sounds, that's and that's what they've sung with for forever. Uh, but the the declarative knowledge is really helping us refine how to guide them to those those sensations. So it's a it's a a a wonderful conversation that um um it needs to be ongoing. So I'm that's why we need to be science-informed voice teachers. Gotcha.
SPEAKER_04Yeah, that's a cool question. Uh we could talk about for another hour, probably. Um, if I were to give a just one response, I I don't want voice teachers to be of afraid of science and afraid of learning more about this, sort of first and foremost. And so this is something I've I've heard Ken say frequently, and I I really believe it too, that um I I don't know a single scientific study that has set out to study what really terrible feeling singing feels like. And or or singing where when you go through a register transition, it is just completely awkward and feels awful. Like, I think most of these phenomena that we can identify, harmonics crossing formants or tuning specific acoustic things in certain ways, or setting up standing waves in certain ways. Like, from the singer's point of view, I I think we are just analyzing after the fact what feels amazing to them. And so I think the information is really useful for teachers as sort of predictive guardrails. Like you really can learn a lot about how a voice, uh, all other things being equal, like should be able to behave. And so you can make predictions about what sounds you should ask of the people that you're working with. But at the end of the day, I I think what I think what all of the knowledge points to is that you know, singing is just a magnificent wild yop that the human body can make, and it feels great, and we should never attempt to um reduce it or sterilize it by imposing language on it that is just is just trying to capture how great singing is.
SPEAKER_01Uh well, I've gotten very interested in uh not only human production, but also the production of uh birds and mammals and uh in general the animal world around us. And I'm so amazed by the fact that all species are tuned so sharply to a very, very small section of the total sound that's available to us. And uh and so for me the beauty is to study the sound in its general term um and uh and then and then figure out what is one individual in in a given species or a cross species um really doing with their sensory information, uh with that little sliver of sound that they can receive. Um and uh yeah, to me uh therefore acoustics itself is not any more important uh than perception, but it perception is every bit as important as acoustics, but we have to understand it across individuals and across species.
SPEAKER_03That's great. Well, wonderful. Well, thanks uh uh everybody. I I uh I really appreciate the the discussion. Um uh and uh I'm gonna be thinking about that wild, magnificent yacht line for a little while there, Ian. Thank you very much.
SPEAKER_04And then we can measure it and understand it.
SPEAKER_03Right, right, right, right, exactly. That's that that's the point, right? So um well, I'll go ahead and uh just close us out here. Um so thank you again, Ingo, uh, Ian, uh Ken for really thoughtful, really engaging in the conversation. Um, this has been Toe-to-Toe with Ingo, a podcast from the National Center for Voice and Speech, where evidence matters. Questions are welcome and disagreement drives discovery. Until next time, I'm Andrew Parker. Keep asking good questions and follow where the evidence leads. We'll see you next time.
SPEAKER_00Must have a code that you can live by. And so become yourself because the past is just a good.