ETH-68: Ethernet Audio Interface for Linux
naturalsystems.io
[6 comments hidden]
Or is there no such synchronization, in which case there would be long-term drift?
[5 comments hidden]
[2 comments hidden]
surely that works only as long as eth68 is the only interface? but often you might want other interfaces too which inherently have their own clocks
[3 comments hidden]
[2 comments hidden]
Milan-AVB is a deterministic, standards-based media networking protocol built on top of Audio Video Bridging (AVB) and Time-Sensitive Networking (TSN) IEEE standards. Maintained by the Avnu Alliance, it is designed for the professional audio, video, and live event industries to guarantee plug-and-play interoperability across devices from different manufacturers
[hidden]
[8 comments hidden]
[3 comments hidden]
I have my eye on some other codecs for the next project.
[2 comments hidden]
[hidden]
[4 comments hidden]
Edit: For the context: https://www.eevblog.com/forum/chat/ti-ne5532-audio-opamp-cha...
[15 comments hidden]
I wonder if gigabit would make a difference on latency, faster packet transmissions. Could packets drop from 64 to 32b?
4ms is pretty good but I feel like sub 2ms would be nicer.
[hidden]
It seems like, at 16000TbaseT anyways, you’re adding an overhead of about 150% on top of copper/fiber latency plus transmit time latency to process data at that bandwidth. I wonder if the same holds true at 100 vs 1000, 2500, 10000? Certainly this is a known tradeoff for DDR performance tuning — if you don’t mind spiking response times greatly, you can get the advertised maximum speeds, else you accept less bandwidth for somewhat less latency — and they’re both effectively using the same strategies to talk over copper.
[7 comments hidden]
[5 comments hidden]
(emphasis mine)
The latency depends on the device. Hardware implementations of Dante commonly support latencies of 1ms or less, but software implementations are higher. The minimum latency of Dante Virtual Soundcard running on a PC is 4ms.
That's the appropriate number to compare against here (since the PC is using a software driver to interface with the network). However, that 4ms number is one-way latency, and the OP's 3.6ms number is round-trip. So this is already half the latency of DVS. (That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.)
[3 comments hidden]
Hm, the audio latency doesn't fluctuate so it's not like the testing conditions affect the measurement. The audio latency is a fixed quantity that depends completely on the number of storage elements in the data path which isn't variable.
Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting. The histograms on my page (which should be animated BTW) show real time processing latency measurements while running the audio all the way through Bitwig with a moderate DSP load (multiple instances of Pianoteq, samplers, live MIDI input). On my system (details at the bottom of my page), I can do this at 48 kHz and 64 sample buffers with zero underruns. If I drop down to 32 samples per buffer, I do start getting underruns.
All I can do from the hardware side is try to minimize the processing latency of a typical cycle so that there is more head room to absorb jitter. The vast majority of the jitter comes from the Linux host. It's up to the end user to tune the system for low jitter. This is usually the case for audio on Linux, and the rabbit hole can go pretty deep on system tuning.
[2 comments hidden]
Right, that's what I meant. There will be some jitter depending on the scheduler, system load, network performance and traffic, and the quality of hardware/drivers; so more latency gives you headroom to absorb the jitter without underruns. The "ideal configurations" I was referring to were the low-traffic network and high-quality NIC. On a setup with more jitter, you might have to increase the buffer size (and thus latency) for reliable operation.
Mostly, I was just trying to contextualize the numbers for readers who aren't super familiar with low-latency audio networking. Sure, this project may not achieve the sub-1-ms roundtrip latencies that you can get with dedicated Dante hardware (like a Yamaha mixer and stagebox); but Dante can't do better than 8ms when one of the ends is a PC (although they were probably aiming for reliability on setups not aggressively tuned for minimum jitter.)
[hidden]
[hidden]
[hidden]
[3 comments hidden]
[2 comments hidden]
If I'm playing on a musical keyboard trying to add a new track while listening to existing ones, the tempo of the two doesn't quite line up and it's jarring. I've had to switch to an old, dedicated PCI audio card.
Would this resolve that?
Being able to use a spare Ethernet jack and throw away the card would be great.
[hidden]
[5 comments hidden]
[3 comments hidden]
[hidden]
ETH-68 doesn't have nearly as broad of scope as AES67. It is a much simpler system, so easier to realize for a 1-person team working on weekends.
Just a few other comments: 1. AES67 usually means Dante hardware which is notoriously expensive. 2. Dante devices running in AES67 mode often (always?) have their capabilities reduced. At least in some cases you are limited to 48 kHz, and I believe you are also limited on latency compensation settings. Someone with more experience could chime in.
[10 comments hidden]
Don't get me wrong, I'm genuinely trying to understand.
How does this compare to pipewire over ethernet? Is it realtime vs buffered?
[6 comments hidden]
Imagine having this on the stage right next to your analog gear, and a computer 50m away.
(Opening the page under discussion was actually helpful, it lists all this right at the top.)
[hidden]
But this might be even more effective.
Besides audio dsp for live use, there is also the use case of visualizers.
[4 comments hidden]
[2 comments hidden]
ETH-68 isn’t running PipeWire, or Linux. Maybe I don’t understand your question.
Not sure what you mean by realtime vs buffered. Care to elaborate?
[hidden]
This enables me the watch video with no latency problems, because it's embedded in the audio stack (buffered). That's why I was wondering which problem gets solved with this project.
But now I understand that this is for concerts not for some audiophile multi room setup.
[hidden]
Pro audio systems frequently (and increasingly) use networked audio. Some obvious uses are distributing the sound from the instruments and people on a stage over to the front-of-house mix position that's usually somewhere mid-crowd, and also to the monitor mix position that's usually in a vaguely-quieter area off to the side of the stage, and to the broadcast truck.
The old tried-and-true method also still works: Analog splits. Take a bunch of audio sources (eg, microphones) and plug them into passive stage boxes that output over a thick-ass cable. Those thick-ass cables go to larger passive split boxes (often on wheels by this point), with two or more outputs for even-thicker cables, with one pair of wires for every individual signal -- often with individual shielding and jacketing.
Eventually, these splits can deliver audio to the different places that need it -- where it's ultimately broken back out into a bazillion individual cables that get plugged into things like mixers.
The cables can be very long (hundreds of meters) in length, and extremely heavy. They're expensive to produce, they're expensive to maintain, and they're expensive to wrangle. They often get transported in their own dedicated wooden trunks. But at least it's simple: A bunch of different audio devices scattered all over a venue, wired in parallel, listening to the signals that are directly produced by microphones on a stage.
---
But with networked audio, it can be more like this: A few boxes on a stage that accept analog audio on one side and emit network frames (often Ethernet or Ethernet-adjacent, and actually using IP isn't a rule at all) that contain digital audio on the other side. Those frames go to a network switch. One or more tiny-ass network cable comes out of the switch and goes wherever it needs to go, and switches can be cascaded, and more audio channels can be added downstream. Because Ethernet(ish) is a many-to-many network, it's bidirectional, too: Audio signals can go upstream just as easily as they go downstream.
It's tidy. It works. It's still expensive because the endpoints are expensive, but the cables themselves can be fairly inexpensive (think robustly-built Cat6 or fiber patch cords instead of giant cable trunks). If the venue's infrastructure goes to the right places and can be trusted, then it can also be used: Plug the stuff from the stage switch into a fiber patch panel on the building, and plug the broadcast truck outside into the same building, tie them together in some MDF or IDF somewhere, and send it. (And in a pure and just world where dedicated fiber links both exist and are easy: Patch another into the studio downtown. Or route it over an IP link that is shared with other purposes, if appropriate and also feeling brave.)
But with the tidiness comes complexity. Like... Latency is kind of a big deal here in ways that aren't a practical issue with analog audio. Putting too much delay between a vocalist and the monitors that they hear themselves with is actively deleterious of their ability to sing, for example.
And buffers are still required (they're ~always required when packet-switched network frames get converted to continuous analog signals). Keeping the buffers small requires very tight timing signals that get shared between all points. Pre-existing systems often achieve that with things like Precision Timing Protocol (though variations exist).
That all conspires to mean that the heavy lifting at the endpoints is often in the realm of FPGAs.
But, again: We get many channels over some bog-standard network cabling. Dante, for example, can be used to transport hundreds of 48KHz 24-bit audio channels on one gigabit ethernet link.
---
Anyway: This is a cheaper, smaller method. It uses an STM32H7 microcontroller to convert betwixt the network transport stuff and the DACs and ADCs of the analog world. It's designed to be used with the open-source Jack system that is commonly-used internally whenever Linux gets involved in recording or stage use, so it's simple to integrate with a Linux PC running software like Reaper. And at the end of the day, it's transportable over the Ethernet networks we all have.
And despite being built around an STM32, it achieves quite usable latency: The stated 3.620ms is about the same as a 1.2 meters of distance for sound in air.
Neat stuff. I'll probably never use it, but it's neat. :)
[22 comments hidden]
[18 comments hidden]
[4 comments hidden]
[3 comments hidden]
There ARE good reasons for recording and processing music at higher bitrates and sample sizes. Effects, especially those with positive feedback loops, can go more unstable and clip and lose information with fewer bits and samples.
There's just no good reason to DISTRIBUTE music at the higher rates.
[2 comments hidden]
[hidden]
Better to upsample once (generally a non-lossy operation) and operate on single precision FP (24-bit mantissa and an exponent) at higher sampling from that point forward and then downsample once (generally a LOSSY operation).
[6 comments hidden]
[4 comments hidden]
[3 comments hidden]
[2 comments hidden]
But maybe there's a use case for replacing an AES192 signal carrying the full FM baseband signal over ETH-67? I'm not sure it's the correct fit, though.
[hidden]
https://www.telosalliance.com/radio-processing/audio-interfa... etc.
[hidden]
A niche use, mind you.
Basically, I found myself doing real-to-complex baseband conversion for audio signals, and doing so efficiently halved my Nyquist rate.
Increasing my sample rate to 96 kHz let me construct the 48 kHz analytic signal that I wanted.
Hoisted on my own petard :)
[2 comments hidden]
[hidden]
I need to record ultrasound for work onboard construction vessels, currently we use either specialized equipment for PAM, which is limited in some aspects, or a USB sound card with a SBC and a hacky setup for sending PCM over TCP.
[4 comments hidden]
[3 comments hidden]
Just curious though, since I don’t work at these sample rates. Is resampling not good enough? Or maybe I don’t understand the mastering flow.
[hidden]
[6 comments hidden]
[5 comments hidden]
[4 comments hidden]
[hidden]
[7 comments hidden]
Um.
[4 comments hidden]
[3 comments hidden]
See, for example, Cloudflare's stories of issues caused by misconfigured systems appropriating the "1.1.1.1" address for their own purposes: https://blog.cloudflare.com/fixing-reachability-to-1-1-1-1-g...
[3 comments hidden]
[2 comments hidden]
https://github.com/gavv/libASPL
I have also thought some about whether or not I could make BlackHole work.
My takeaway so far is that macOS is more of a hassle for developing this stuff. It doesn't help that CoreAudio is not open source and the docs are hard to navigate. Hopefully someone will fix JACK-router on macOS one of these days.
[hidden]
[8 comments hidden]
"very low latency" in audio is <=1ms. 3.6ms is good but not special.
[2 comments hidden]
> The LATMON pulse width is therefore an accurate measure of the total processing latency of each cycle and is affected by every element in the data path: processing delay in the microcontroller, network transmission delay, host OS delays, signal processing delay in the DAW, etc
[hidden]
One way audio latency is about 3.6 ms divided by two = 1.8 ms. This type of latency is audible.
Processing latency is just the amount of time it takes to complete all processing for each audio cycle. From the histograms, the typical processing latency is about 625 us and of course there is some jitter (almost all of the jitter comes from the Linux host BTW). Processing latency is not audible. However if the processing latency exceeds the deadline on a given audio cycle, there will be an underrun which will cause an audible glitch.
[3 comments hidden]
I would love to know why this is hand-wavy and not very cool? Even as not-audiophile, I’ve got ideas of things to use this for.
> RME HDSPe AIO Pro PCIe
> eth68 matches the latency performance of the RME card at 48 kHz and surpasses it by 0.33 ms at 96 kHz.
[2 comments hidden]
On Linux, we don’t have much in the way of Thunderbolt support for audio interfaces, so the only way to achieve this sort of latency has traditionally been with PCIe or PCI audio interfaces. Having a low-cost, infinitely more portable solution would be very welcome.
[hidden]
You might be thinking of latency of the converters, which is normally sub-ms.
[hidden]
https://interfacinglinux.com/linux-compatible-audio-interfac...
AFAIK the best one is RME AIO Pro PCIe card. ETH-68 matches the round trip audio latency of this card at 48 kHz and surpasses it be 0.33 ms at 96 kHz.
alowell[34 comments hidden]
iFreilicht[4 comments hidden]
alowell[hidden]
yellowapple[hidden]
4dregress[hidden]
zbrozek[2 comments hidden]
alowell[hidden]
hommelix[2 comments hidden]
alowell[hidden]
NewJazz[2 comments hidden]
alowell[hidden]
herczegzsolt[5 comments hidden]
olpad[4 comments hidden]
NewJazz[2 comments hidden]
olpad[hidden]
alowell[hidden]
Gracana[9 comments hidden]
https://radar.cloudflare.com/domains/feedback/naturalsystems...
alfanick[8 comments hidden]
Gracana[2 comments hidden]
I went ahead and chose "blog" because it sounds like it's a blog. Apparently the "newly seen domains" category goes away on its own after 30 days, so that would have fixed itself, though then it would have been "uncategorized" and still blocked.
monster_truck[hidden]
dspillett[5 comments hidden]
Yes. And no. The affected users usually have no control over the situation, so if the site owner cares at all that those people can't access it, then it is effectively their problem as they can potentially do something about it.
If neither party cares enough, then it is nobody's problem.
ciupicri[2 comments hidden]
P.S. By the way, to paraphrase an OS developer, Cloudflare, fuck you for "Sorry, you have been blocked. You are unable to access thunderbird.net".
dspillett[hidden]
Often not, because as I said in what you replied to:
> The affected users usually have no control over the situation
The affected users are usually on devices controlled by others or on their own devices connected through networks controlled by others: at work, in schools, in libraries, etc. The affected users and the those configuring the "new site" block (or leaving it on by default) are not the same set in many cases.
My work users such a filter, there is nothing I can do on their laptop to get around it¹, though my phone newer connects to their network so I can just look at things that way if I care to.
--------
[1] OK, there are things I could do, but they would be an abuse of the admin level access I have on the machine, because some things needed for my job don't work properly without that, and it isn't worth that arguement if noticed!
Gracana[2 comments hidden]
I guess some people thought I liked cloudflare or something... that's definitely not the case.
noAnswer[hidden]
The only real solution, as a "holder of a young domain" is to wait 30 days before you use it.
olpad[3 comments hidden]
alowell[2 comments hidden]
olpad[hidden]
vegadw[3 comments hidden]
jaen[hidden]
NewJazz[hidden]
hdb2[hidden]
tech nerd/musician here, I would be very interested in this! do you have an email list or something I can sign up for? this is fantastic, and if the $$$ is not crazy, I would want this for sure.
russdill[2 comments hidden]
alowell[hidden]