Blog_1

1. From Physical Sound Superposition to a New Direction

At the beginning, the project had a different shape in my mind. Last semester I was mainly interested in physical sound superposition: real loudspeakers in a room, sine tones playing at close frequencies, and the listener walking through zones where the sound changes. I liked the idea that the room itself could become a kind of instrument. The piece was not only about listening to tones, but about moving through them and discovering how interference, beating and phase relationships appear in space.

That first idea was more installation-based. I imagined loudspeakers placed in a room and the listener moving physically between them. The focus was on the direct experience of sound in space, not on a screen or an interface. This still feels important to me. Even now, after the project has changed, the original interest is still there: movement, sine tones, spatial relationships, and the strange physical feeling that happens when simple sounds interact.

But as the semester continued, the project moved away from being only a physical loudspeaker installation. In meetings with my professor, it became clear that I needed a more controlled research step before going into the real room. If I directly built a physical installation, many things would happen at once: reflections, room modes, loudspeaker differences, directivity, occlusion, and the listener’s own perception. That can be interesting artistically, but it is difficult to study clearly.

This is where the idea of auralisation entered the project. Instead of trying to simulate the full room, the new direction became: build a controlled direct-sound reference first. In this version, I can define source positions, listener position, distance, direction, level and delay. Then later I can compare that controlled version with the physical room.

So the project did not completely abandon the first idea. It changed its method. The original project was about sound interaction in a real space. The current project builds a headphone-based interactive prototype that can later be tested against a real multi-loudspeaker room. I think this change made the project more realistic as a thesis. It still comes from the same curiosity, but it now has a clearer technical and research structure.

The main question also became more careful. I am no longer asking whether I can completely simulate a room. That would be too large and too easy to overclaim. Instead, the current question is closer to this: how can an interactive binaural auralisation help compare ideal direct-sound behaviour with a real multi-loudspeaker room?

That change in wording matters. It gives the project a better foundation. I can say what the system does, but also what it does not do yet.

Final thoughts and next steps | R&D 2 | Blog 8

Overall I must say, it was quite a journey. Navigating my way through the various errors and setbacks I faced was truly soul crushing at times. I often didn’t know how to proceed and felt completely frozen and helpless. However, I always received assistance from my mentor, Mr. Sontacchi, who motivated me to push through till the end. He has guided me through many obstacles I faced and helped me re-establish interest in the project when I thought I couldn’t make it happen. I would like to personally thank him for all the effort he has put in and the support he’s given me throughout this semester.

One thing I wish I had was more time to focus specifically on this project. There were various scenarios where I was caught between multiple deadlines, which restricted me from focusing on what I wanted to achieve. It was frustrating as I wanted to progress and move forward with the work but was held back due to the workload from the other courses. Balancing the time for this research project alongside the other project felt difficult. I felt this project could have been even more expansive if I had more time to contribute to it. 

One of the things that I had to change was the model of the waveform for the visualizer. I initially intended to work on a 3D waveform, with more particle clouds and points, but it turned out to be more complicated than I had initially expected. Through my mentor’s guidance, I built the existing model, which varies in pitch based on its position on the y-axis. I really like the current model, but would like to rework it and build an alternative for the final installation. 

Additionally, I wish to integrate the pitch to colour mapping in some form. I am planning to map harmonics and timbre to the properties of the newly designed visualizer. I also need to implement a surround sound system for the audio response to recreate an immersive environment, as well as offer additional control parameters such as playback loudness, directional panning and wet/dry mix of the output. There is a lot of testing that needs to be done to ensure we don’t suffer from latency issues.

In the end I am happy with the output and super glad to have a functioning prototype to present. That being said, I still believe there is a lot of work to be done in the upcoming semester for me to completely realize the installation that I had envisioned. I am looking forward to continuing work on this research project.


Functional Prototype with added features | R&D 2 | Blog 7

My prototype for this semester is finally complete. After much trial and error, I was able to build a visualizer that functions as per the user’s input and is represented in a model that is easily understandable to the user. The waveform morphs and reacts based on the user input and doesn’t break or alter its core geometry in the process, finally maintaining a smooth and fluid visual experience. There are still a few additional tweaks that can be done, but it’s somewhat functional at the moment, which means that I have met my primary goals for this semester.

Some additions include the audio feedback response, which acts as a monitor based on chords to determine the pitch accuracy of the user. This was done using the quality parameter in the sigmund object, a very useful tool for pitch tracking. The signal runs through a relational operator (>) and when the signal fulfills the condition (greater than 90%), the patch plays back a chord for the user from an oscillator based on their pitch quality. If the user is perfectly in pitch, it will play a pleasant major chord, while a slightly off-pitch input would trigger a more dissonant minor chord, providing immediate auditory feedback without the user needing to look at a screen.

Another addition is the waveform, which now features a visual activation response where the visualiser remains inactive or dark when there is little to no input signal. The colour in the waveform reappears when the user is speaking, and this is a vital visual feature to indicate that the signal is being successfully received from the user. By fading to dark during silence, the interface reduces visual fatigue and makes the interactive moments feel much more dynamic. This addition was done using the data from the quality argument in the sigmund object, which was mapped to a Logic CHOP in TouchDesigner to trigger the visual state changes seamlessly.

I also considered adding a few advanced audio analysis tools from a Python library to give users more speech information. What I had imagined was an algorithm which collects a live data feed and predicts the different groups of consonants present in the speech (fricatives, affricates, plosives, etc). You would see numbers fluctuating on screen showing the variations in frequency as well as the tonal quality.  Unfortunately, I won’t have enough time to implement this into the prototype, and it would be considered as a feature to be added in the future.

In the final blog, I will share my overall experience building the prototype, reflect on the challenges, and give insights on what to look forward to in the upcoming semester.

Ars Electronica Center Blog Post

Last month, the sound designers and exhibition designers of FH Joanneum jointly participated in an excursion to the Ars Electronica Center in Linz. The center is a museum that focuses on new media art and is most renowned for its annual festival, the Ars Electronica Festival, which critics have praised as a pivotal exhibition in the digital arts. It is widely considered one of the best known creative arts festivals in Austria. Over the years, the people behind the center have been praised for their unique integration of technology into the creative field. They also have a dedicated inhouse research facility called the Futurelab, which focuses on modern technological advancements, particularly artificial intelligence. 

There were a lot of engaging exhibits. Some memorable ones included the fascinating “tardigrade” under a microscope, as it made its way through its tiny surroundings. These remarkable creatures are incredibly resilient, capable of withstanding extreme temperatures and are known to be some of the only organisms able to survive in outer space. Another interesting part of the excursion was the Deep Space 8K experience, which drew us into a world of 3D visuals. We heard various soundscapes, ranging from the planets of our solar system to our oceans, and even the unique art of yodelling.

The exhibit I would like to speak about in detail is the simple sequencer, which was located in the Kids Research Lab on the first floor of the building. I really enjoyed the sequencer and felt like a kid again while playing around with it for some time. It was a really nice setup, especially for children who are curious about music and sound in general. It was a simple concept, but one that was very much fun to use due to the various patterns you could recreate.

The sequencer was off by default, so you had to place blocks (or tiles) representing a particular instrument onto the vacant slots to produce sounds. On the left hand side of the instrument block was the fill pattern, consisting of 8 distinct beats that the user could customize to recreate a unique rhythm. There was an LED strip on the top edge of the sequencer, which indicated the beat being played at the time. The beat blocks came in different colours and sizes. Some were taller than the others and had unique shapes on the top facing size. However, I was not able to decipher exactly what these could have stood for. This may have possibly been a measure used for quarter or half notes, but I am not entirely sure. The instrument blocks included an acoustic guitar, a piano, an electric bass and even a saxophone.

The sequencer could only accommodate 4 instruments at a time as there were only 4 slots available. Apart from this, there was also a small blue box on the right side of the sequencer with buttons featuring images of snares, cymbals and kick drums. I do not recall if these sounds played in sync with the sequencer but pressing them did produce the sounds of the instruments I just described.

Overall, I really enjoyed the visit and gained valuable insights from the unique exhibits. From the way they were presented to the overall setup of the installations, it was a good trip to explore and brainstorm ideas for our own installations at Klanglicht later this year. 

Errors, Fixes & Reiterations | R&D 2 | Blog 6

Building the visualizer was the most challenging part of the project so far. I ran into a lot of trouble fixing the waveform as it kept changing drastically through each iteration. In this post, I will give a breakdown of the workflow and briefly discuss the multiple errors I faced along the way.

I built the ribbon based visual system using multiple elements as I mentioned in the last blog, utilizing various combinations of multiple CHOPs (channel operators), TOPs (texture operators) and SOPs (surface operators) to get the output. I initially used an OSC In CHOP to collect the data on pitch and loudness. These data streams were then separated using a select CHOP and sent to a math CHOP to define the parameter ranges. This was a crucial step, as defining these ranges significantly altered the visualizer’s behaviour. I finally settled on a ‘from range’ of 40 to 100 and a ‘to range’ of 0 to 4 for the pitch. The loudness, meanwhile, has a ‘from range’ of 25 to 100 and the ‘to range’ of 0 to 2.

I had to create more math CHOPs to carry out more calculations. I split the loudness in half and combined it with the pitch values through the operation of addition and subtraction. This was done to define the upper and lower limit of the waveform, which were called ‘ty_top’ and ‘ty_bottom’ respectively. After this series of math CHOPs I had to place a Trail CHOP to capture the history of the data, recreating the continuous stream of data at the given timeframe. This was followed by the pattern CHOP, which served as data for the x-axis. The pattern chop was then sent through a math chop and merged with the individual channels (ty_top and ty_bottom). Once this was done, we converted the CHOP to SOP for adding surfaces and textures. Finally, the SOP was placed into a Geometry COMP then a camera TOP and a render TOP were added to the end of the signal chain to generate the final visual output.

This basically makes up the entire visual build, but there were a lot of errors that occurred along the way. The first reoccuring issue was the appearance of a static line despite signals being received through OSC. Another visual bug was these sharp, triangular shapes that kept being pushed from the center, it looked very confusing. I also had multiple scenarios where the signal did not rise along the y-axis with the variation in pitch. And at times, we also had signals literally break in two halfway through the visualizer. It became so bad I started a new visual project in Processing, as it felt more likely to give me the output I was expecting.

In the end, I am trying to achieve a functional model with a few additional features that help the visualizer feel more interactive and complete.

Deciphering OSC data & Visual Model | R&D 2 | Blog 5

In the last blog, I discussed the functioning and features of the audio pipeline, which was built using plugdata. We were able to collect pitch and loudness data from our audio stream and were attempting to send it to TouchDesigner via OSC. However, the connection was not successful. I had to redo the patch and will share my insights and methodology on fixing the issues.

I had to repackage the data streams into an [oscformat] object and then prepend them into a list before moving on to the next step, which was to replace the [oscsend] object I had previously used with a [netsend] object containing the arguments ‘-u’ and ‘-b’. This sends data through the UDP protocol. It was configured to send messages to localhost on the port I set at 3000. 

In TouchDesigner, I used the ‘OSC In CHOP’ to collect the information sent from plugdata. I had a small monitor indicating the variations in values inside the CHOP, this meant that the data connection was successful. I was now ready to build the audio reactive visualizer. I wanted to see how the OSC data affected the visual output. So, I tried testing it out with a project I found through the Youtube channel supermarket sallad. It was a spherical visualizer which simulated various particles and noise. It was a visually appealing piece, I was able to play around with the lighting along with particles and found it really interesting. I wanted to implement a similar visual model with particle clouds for my project but later decided to go with an alternate approach.

My initial concept was to build a 3D waveform model for the visualizer, but it was proving to be quite difficult. After consulting with my mentor, he suggested that I work on a simpler model and if needed modify it later on. I was now trying to think and come up with ideas for new visualizer styles which can also be easily understood by a user. So, I kept searching and eventually got curious with the waveform structures found in pitch shifting plugins like Melodyne. It seemed like an interesting model as you could tell the variation in pitch based on different positions along the y-axis. This also made me move away from mapping pitches to colour (Camelot wheel) as I had described earlier.

I would still have to consider a couple of things to see if this model works for me. What kind of elements (CHOPs, SOPs, TOPs) would I require to build this model and how accurately can I display the concept I had initially planned? Some of these will be answered in my next blog. Until then!

Pix2Pix: GANgadse

When I stood in front of the interactive screen at the Ars Electronica Center, there was one thing that sparked my interest immediately. The installation invites you to do something incredibly simple: pick up a digital pen, sketch a few lines, and watch a machine learning system immediately try to transform your doodles into a fully rendered, colorful cat. The piece plays on a popular piece of German internet slang for a cat, setting a lighthearted tone for what is actually an existential encounter with artificial intelligence. I drew a few shaky, anatomically questionable lines, and almost instantly, a furry, slightly cursed digital creature arose on the screen.

Beneath the humor of creating these accidental monsters lies a fascinating look into how modern generative AI interprets human input. The engine driving this transformation is an advanced neural network known as a Conditional Generative Adversarial Network. To understand how it brought my terrible drawings to life, I found it helpful to picture a high-stakes creative competition happening inside the computer. The system splits into two competing algorithms: a Generator and a Discriminator. The Generator acts like an art forger, starting with no knowledge of what a cat looks like and trying to create one from scratch. The Discriminator acts as a detective, comparing the forger’s creations against thousands of real cat photos it memorized during training. They push each other until the fake images become astonishingly detailed.

What makes this specific setup so fascinating to interact with is the conditional part of the tech, which is designed for image-to-image translation. In a standard setup, you press a button, and the AI spits out a random, perfect image. Here, the system is given a strict blueprint: a simple sketch. The Generator is forced to translate my exact lines, curves, and mistakes into the final image. It looks at the brushstrokes and figures out how to cram the textures of fur, whiskers, and shadows into the bizarre boundaries provided.

This translation process relies on the network’s ability to recognize spatial patterns. For my first attempt, captured in the image above, I tried to play along and drew a relatively standard, cartoonish cat face with big round eyes and pointed ears. Because the drawing roughly aligned with what the AI expected, it tried its best to map realistic textures over those specific regions. However, you can see how it over-interpreted the massive eyes I drew, filling them with an unsettlingly realistic, glossy depth that makes the final output look incredibly intense, yet undeniably feline.

Because this network was trained exclusively on felines, it possesses a hilarious, stubborn blindness. It is completely incapable of seeing anything else. If you try to draw a house, a car, or something else entirely, the system will still desperately search your lines for pointy ears or whiskers, forcing cat attributes onto absolutely everything. I decided to test the absolute limits of this bias with my next drawing, which you can see in the image below. Instead of a cat, I drew a stylized character wearing a backwards cap, sporting giant elephant-like ears, and a long trunk-like shape on its face. The AI was completely unfazed by my lack of cooperation. It looked at the round head and the brim of the hat and somehow translated those shapes into a furry, shadowy texture, attempting to force the contours of an animal coat onto human streetwear.

This rigid worldview is exactly why the final drawings turn out so beautifully bizarre. The machine has no conceptual understanding of biology, anatomy, clothing, or what a living creature actually is; it only understands pixel statistics. When I drew an impossible, abstract shape, the network faithfully attempted to render photorealistic fur, depth, and organic lighting over my nonsensical geometry. The result was a surreal hybrid—a digital creature that looked like a cubist painting brought to life with organic textures.

Walking away from the screen, I realized the installation is a brilliant educational tool wrapped in a playful artistic experience. It peels back the layers of the mysterious AI black box and lets you see firsthand how these algorithms interpret, distort, and reconstruct our world. It left me thinking about a future where human-machine collaboration looks exactly like this: we provide the messy, creative spark through a simple sketch, and the machine handles the complex, data-driven task of rendering it into reality.

Building the Prototype | R&D 2 | Blog 4

The first step to build my working prototype was to start with the Pure Data patch. This is an integral part of the project as all the live data is being collected and parsed through these connections. I used an alternate version of Pure data called plugdata. I found it more convenient to work with as it has a much more modern UI, helping users (especially me) navigate the software easily.

So, my intention behind this patch was to build a system to collect live audio signals, retrieve information about pitch and loudness of these signals, and store them in lists which can be sent to OSC for TouchDesigner to collect and interpret. The audio signal is captured using an [adc~] object which is then amplified using a [*~] object which is connected to a horizontal slider to ensure input gain can be controlled by the user. This is one of our most important control parameters. The patch then analyzes the audio stream and splits it into the two properties that I mentioned earlier, pitch and loudness.

In order to obtain loudness information, we had to use the RMS value which can be computed using the [env~] object. I wanted to use decibels for the project, as it is the most commonly used measurement for playback systems. In order to achieve this, I used the [rmstodb] object. To test if the signal was functioning as expected, a VU meter was connected to monitor the signal. This was then connected to an oscsend for the data to be transmitted.

As for the pitch, it was slightly more complex. I had to use the [sigmund~] object with the arguments ‘-pitch’ and ‘-peaks’ to obtain the values. I was not familiar with this object and it took me some time to understand the working behind it. It continuously analyzes the waveform to track fundamental frequency. The frequency value is then passed through an [ftom] object which converts frequency into MIDI objects. This helps to work with musical notes that are more familiar. It is also much easier to map to visual parameters. These MIDI values are also sent through OSC for TouchDesigner to collect and process.

My next step is to test the OSC data and ensure it is being received correctly by TouchDesigner. Should be simpler considering the data being transferred locally. Localhost should suffice for this and netsend might also do the trick. My findings will be revealed in the next blog. Stay tuned!

Defining Parameters and Mapping | R&D 2 | Blog 3

One thing I had to take into consideration while building the prototype was defining the parameters we were going to use for the project. Obviously, being an audio reactive visualizer, the two most relevant (and important) parameters to map would be the loudness and pitch of the input signal produced by the user. These quantities will be mapped to affect the position, shape and texture of the visualizer. Sounds simple enough. So, how exactly will this be done? What is affecting what here? To answer that question, we need to delve a little bit into colour theory based on the Camelot wheel.

My idea for mapping the user’s pitch is to base it on the Camelot wheel, a visual tool often used by DJs to help with harmonic mixing. It is based on the principle of the ‘Circle of Fifths’, which maps out all 12 musical keys in a visual clock format, highlighting the different relations between them. The Camelot wheel turns this into a shorthand guide by using numbers (1-12) denoting each fundamental musical key and letters (A & B) representing the scale types. ‘A’ indicates minor keys while ‘B’ highlights major keys. It is usually visualized as a rainbow wheel, but it does not really hold any true correlation between the colours and the notes. I would like to build a type of cross modal correspondence and relate colours to notes, indicating variations in pitch through colour differences.

The loudness parameter was a more intuitive choice. I based this on the waveforms we see today, a very obvious correlation to represent loudness levels. The louder the sound, the higher the peak of the waveform. It will be interesting to see if we could recreate a highly dynamic waveform based on the user’s voice. This also serves as a commentary (no pun intended) on the loudness wars, which we still see today. Most tracks produced today are so heavily compressed to the point that they appear as continuous peak waveforms, losing the unique dynamic quality of each particular instrument and audiotrack.

Two other things we need to consider are the sensitivity of the visualizer’s response to the input and how it can be fine tuned to get the best possible results. We also need to ensure the dataset is normalized such that it is suitable for our visualizer to interpret. Because huge variations in range can massively alter and shift the output. The biggest hurdle will be finding the exact range or measure to ensure the visualizer functions exactly as we need it to.