Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> There's no such thing as a codec that looks good or bad in an absolute sense.

There are plenty of examples of both awful codecs (software; encoders) and of digital audio/video standards (e.g. Vorbis and MPEG-4 ASP) suffering from material limitations.

> If you gave DVD levels of bits-per-pixel to MPEG-4 ASP you could get something that looks nearly perfect.

No. Not even if the source is encoded by XviD. Not all problems can be solved by throwing higher bitrate at it. The ASP only supports 4:2:0 chroma subsampling, which is the largest of several shortcomings contributing to the limited quality you can achieve with ASP video material.

People on at least an intermediate knowledge level of digital video know what a damning problem 4:2:0 subsampling is and how heavy a penalty it incurs on color and clarity. Your comments mostly hold a suggestion that you kinda have some rough idea about digital video. And that's OK.

MPEG-4 SSTP is a very different matter, but that's not what DivX and XviD encodes.



I should have added "general purpose". I know there are restrained encoders for specific situations. Can you list a few of your examples for video? And I don't understand the other example you did mention, isn't vorbis more than enough at high bitrate?

4:2:0 is just fine for video. It is not a heavy penalty. It's OK for you to be rough about this. More seriously, if it's good enough for bluray and UHD bluray then it's fine.


Yeah sure, a few example that spring to mind would be...

Poor encoder: FAAC/FAAC2, the first open-source encoder for AAC audio, produce terrible audio no matter how much bitrate you let it work with. The AAC standard itself facilitates crisp audio quality at low bitrates, as heard with e.g. Apple's Core Audio AAC encoder or Nero AAC.

Poor encoder: Xing, a popular MP3 encoder of the early 2000s, was similarly infamous for producing chirpy and slurry audio even at or above 192 kbps, while bona fide MP3 encoders like LAME do far better on less.

Poor encoder: NVENC, Nvidia's on-GPU hardware video encoder, produce very poor H.264 video even at 8-10 mbps, even on the current 8th and 9th generation (RTX 40/50 series). Good H.264 encoders like x264 is capable of producing excellent FullHD video at just 2-3 mbps.

Poor standard: Vorbis is a good example of a spec whose limits/mistakes make it impossible to preserve certain combinations of frequencies, resulting in brief passages where parts of the reproduced spectrum deflates, making some music sound as if it lost its breath, so to speak. When fed certain "triggering" audio content designed to expose problems in the spec, Xiph's reference Vorbis encoder will produce ringing sounds. Interestingly also the MP3 spec has similar limitations where certain frequency combinations (usually towards the lower and upper ends) will reproduce with quantized amplitude, even when encoded with LAME, though the outcome is nowhere near as pronounced as it can be with Vorbis.

Poor standard: MPEG-4 ASP, being limited to 4:2:0 chroma subsampling and PAL/NTSC resolutions. My beef with 4:2:0 is because of how harsh it is on low-resolution content. MPEG-4 ASP being limited to a maximum of 720x576, and the 4:2:0 chroma coverage being only a quarter of that, is the reason why DivX/XviD content is smudgy even with reproduction filters.

On FullHD content a 4:2:0 grid has almost three times higher resolution, which I agree works out on both still scenes and slow panning (the two scenarios where low chroma resolution makes itself most reminded).


Oh I didn't mean to waste your time on poor encoders. I'm sure there's many of those. I mean actual codecs that have a problem, that make it impossible to do good quality.

> Vorbis is a good example of a spec whose limits/mistakes make it impossible to preserve certain combinations of frequencies, resulting in brief passages where parts of the reproduced spectrum deflates, making some music sound as if it lost its breath, so to speak.

So the people that talk about bitrates where it's transparent are basically delusional?

> Poor standard: MPEG-4 ASP, being limited to 4:2:0 chroma subsampling and PAL/NTSC resolutions.

That makes a lot of sense, I had no idea it was limited to those resolutions.

Thanks for the time explaining those.


> So the people that talk about bitrates where it's transparent are basically delusional?

In my opinion they are not. A lot of people won't pick up on differences unless something is pretty off, even if they were intimately familiar with the audio beforehand. Some people are barely able to tell the difference between one and the same piece of music being played to them first in stereo and then in mono. A fundamental problem with blind listening tests, such as those often cited from the venerable Hydrogen Audio forums, is that they provide subjective instead of objective truths - besides everyone's hearing being different, an individual's perception of music also changes by the day as it's affected by their current mood, state of health, emotional state, whether they are rested or not, and so on. Applying a scientific method (e.g. PSNR or spectral analysis) reveals an objective truth of how close a lossy product is to the original, which is where Vorbis has been shown to fall short. But subjectively a lot of users may never notice, nor even care.


> A lot of people won't pick up on differences unless something is pretty off, even if they were intimately familiar with the audio beforehand. Some people are barely able to tell the difference between one and the same piece of music being played to them first in stereo and then in mono.

Well anyone that declares something is transparent without "to me" attached, based on only their own testing, is in the flippant "delusional" category.

I'm talking about people that get the best listeners they can find to seek out differences, and combine their knowledge together while hunting for flaws. If they can't even find one or two people that can pick out a distortion, then that's a pretty trustworthy standard, even if that only gives you a 99.9th or 99.99th percentile listener. Are these people aware of the specific problem you're stating? Do they agree or disagree?

> Applying a scientific method (e.g. PSNR or spectral analysis)

But that kind of truth is far from what you actually want, a measurement of how imperfect the compression is to human ears. You can use a complicated mathematical model for how listening works, but you're stuck calibrating that model with humans, punting the problem up a level. Bulk semi-subjective data is the best foundation we can get our hands on.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: