By "small", it's meant in a relative sense. JXL's coded blocks can be as big as 64x64, and block boundaries can create visible seams. Example: https://juliobbv.com/pics/photo.jxl
These seams are especially noticeable in the background, and align with coded block edges. You can tell the encoder is trying to conceal them as best as it can, but this cannot be properly mitigated without a proper deblocking filter.
For comparison, AV1 tiles normatively go through the deblocking filter.
- Upload the problematic image to the AOM analyzer
- Press 'L' to show tiles view
- How many tiles (yellow rectangles) do you count?
- Are artifacts actually at those tile borders?
Worth mentioning: their demo only uses two passes for progressive AVIF -- it's just their choice. You can have up to two more intermediate passes, so at 30% you can have something much closer to the original image.
Even still, the demo proves that AVIF can deliver a usable image with up to 3x as fewer bytes as JXL!
> their demo only uses two passes for progressive AVIF -- it's just their choice
Also worth mentioning: AFAIK "their choice" here is just the default, i.e. what `avifenc --progressive` outputs, so maybe if the default is suboptimal, it could be improved?
> Even still, the demo proves that AVIF can deliver a usable image with up to 3x as fewer bytes as JXL!
Usable for what? As a clearly loading image / placeholder, I personally like the JXL version better (as already indicated in my previous comment). As a final image I think both of them are unusable, and if that's the intention I think it would be better to encode both aiming for very low quality, maybe also reducing the resolution, without progressive loading, then compare. My understanding is that AVIF is usually better at very low bitrates, so it would probably be better, but I don't think this example "proves" that.
Well, the default in avifenc can always be changed. Do keep in mind there's no "one size fits all" implementation, as customers desire different loading tradeoffs. You might be surprised, but during testing (outside HN), we've seen people actually prefer "2 layer" loading.
I'm surprised HN likes progressive loading to be more granular, and use that to push back. I'm wondering if there are generational differences at play? Maybe it's the difference of being used to the blurhash vs. old-school JPEG loading experience.
Anyway, the more expressive mode in avifenc is `--layered` (yes, I know the name is weird).
> Usable for what?
Well, focusing on the "Poke bowl" example: with AVIF you can clearly tell apart each ingredient at 8kB. You can derive enough semantic understanding just from this base layer. Having decent edge preservation helps significantly here.
On the other hand, JXL is just too blurry at 8kB to make sense of the image -- only the egg and carrot are recognizable, maaybe the cucumber? IMO JXL needs the pass at ~28kB to make everything salient, including sprouts and beet. Yes, I know there are subjective effects at play and I'm sure we'll disagree on exact image thresholds, but recognizing objects within an image is so important in real-life use cases.
> Well, the default in avifenc can always be changed.
Sure, that was kind of my point, but I think it's important to point out it's the default. I had read somewhere else in the comments that they just "happened" to use 2 passes (that I now realized was also written by you), that I interpreted as them maybe making a bad decision in the comparison page, but I think it's very reasonable to use the default. I also usually prefer to use the default unless I have a good reason not to do so.
And I think that `--layered` is fine as a name, and I had realized that as the docs were right below the docs for `--progressive`, which do mention it encodes a "layered image", although it could use an example on how to use it effectively.
> I'm surprised HN likes progressive loading to be more granular, and use that to push back.
Not necessarily? Using more passes would probably help in specific the metric in the comment I was responding to, by making AVIF possibly look better on a larger fraction of the decoding time. I mostly interpreted your comment as saying that it would have been better if they used more passes, and that's on me.
The reason I was comparing the breakpoints in the decode time was because I was responding to an allegation that "even there, AVIF looks better for 90% of the decode time", and I decided to point out that's not not the case and chose some pretty clear breakpoints that are very hard to argue about. Very subjectively the actual breakpoints could be a lot earlier, for the Poke the AVIF only looks better (in the sense that you can figure out what's there better) in a 6% range (from 2 to 8%) or maybe 9% (from 2% to 11%) of the decode time. There are many (relatively big) seeds that are just not there on the AVIF side, but you can distinguish on the JXL side, even if they're blurry. Percentages used in reference to the JXL side.
> I'm wondering if there are generational differences at play? Maybe it's the difference of being used to the blurhash vs. old-school JPEG loading experience.
Hmmm I don't know, maybe. TBF with today's usual internet speeds progressive loading isn't as important as it was as more steps make less sense if you have a good connection.
> Well, focusing on the "Poke bowl" example: with AVIF you can clearly tell apart each ingredient at 8kB. You can derive enough semantic understanding just from this base layer. Having decent edge preservation helps significantly here.
Sure, I see what you mean, you can distinguish more image element in the AVIF example, that's true. I don't like the lack of consistency and some of the artifacts to the point where I'd prefer not using progressive encoding, so that's not what I'd call "usable", but that seems like a matter of taste. It surely does look closer to a "final image" than the JXL at that point.
As a note, I think what I'd prefer would be 2 passes, with the 1st pass being around how JXL looks at around 11% / 39 kB. You can get a rough idea of what's the image about and the elements on it, with roughly the right colors, but it's still clearly loading, so there's no confusion, and without many "artifacts" like the progressive AVIF in the example (no idea if AVIF is able to make progressive encoding in a way that's more "blurry"). The caveat is that more steps are better when it takes too long (more than a few seconds) so the user knows it's not stuck.
> Sure, that was kind of my point, but I think it's important to point out it's the default. I had read somewhere else in the comments that they just "happened" to use 2 passes (that I now realized was also written by you), that I interpreted as them maybe making a bad decision in the comparison page, but I think it's very reasonable to use the default.
Yeah, I can see that. By "happened" I meant that the default is good and serves a common use case, but also it can't be expected for a 2-pass AVIF to remotely match JXL's finely-incremental experience. The tricky thing is coming up with a good-enough compromise -- one extreme wants their first pass be more like a blurhash (quality 0, 1/8 scaling), while the other wants a medium quality image (quality 30-40, full scaling), and everybody else is in between.
For reference, `avifenc --progressive` is currently quality 10, 1/2 scaling.
> Hmmm I don't know, maybe. TBF with today's usual internet speeds progressive loading isn't as important as it was as more steps make less sense if you have a good connection.
Interestingly enough, I frequently get reminded of spotty internet connections -- turns out you just need to take the subway hah. This is why I'm so passionate about progressive image loading in general.
> As a note, I think what I'd prefer would be 2 passes, with the 1st pass being around how JXL looks at around 11% / 39 kB.
That sounds like you want your first pass be half scaling, around quality 25:
`image1.png` and `image2.png` can be the same source image. You can blur `image1.png` or even better: add a tiny "loading" icon to indicate the image is still downloading.
Oh, i left the wrong quotation whereas i intended to reply to another message. Throughout the blogpost you're quality matching and comparing encoders based on objective metrics whereas it'd be more telling to get the crowd subjective comparisons. I think it's pretty evident that most codecs benchmaxx to the point of objective metrics being useless.
Ah, no worries! I can't speak for Iris and Aperture, but both "tune IQ" modes in SVT-AV1 and libaom had extensive human evaluations to make sure they weren't accidentally being benchmaxxed at the expense of subjective quality.
It sounds like you don't recognize the relationship between text and subtext.
The entire post amounts to saying "JXL in its first iterations was better than AVIF's more advanced state of development and support, but AVIF very recently got better only after several years of widespread adoption, and I don't think JXL would too".
The arguments given against JXL cannot be reconciled with the history of codec development.
Maybe you're just bad at reading text. Again, read the post carefully. Look around... nobody else had the same interpretation as you, so maybe that should clue you in something might be off with yours.
The post literally goes over known ways that JXL can be improved. The actual argument (and this is stated!) is whether those improvements will be enough to make JXL compelling to use for the web vs. other established formats (quality efficiency, encode/decode time, progressive loading, etc).
Also, don't assume AVIF or even WebP have maxxed out yet :) There are known ways to improve those two too!
I'll tell you what. In good faith, you should ask your friend to fix the numbers in the post specifically in relation to concerns that people have with the testing methodology in relation to threads and the speedtest flag instead of handwaving them and to be forthright about how significant codec optimization historically comes after adoption.
Because it's understandable that you'd be defensive in the comments here, but you're not being honest about what's written.
> The post literally goes over known ways that JXL can be improved.
You think so? Let's name them, yes? Ok, I'll go first:
* Splines. Explicitly derided as "vastly more difficult" and "I have no reason to believe they'd be better anyway".
Now you point out the next one.
You see, any unfilled area is specifically mentioned only to dismiss it as meaningless. Meanwhile, out of the other side of your mouth, you say "don't assume AVIF or even WebP have maxxed out yet". Sure. I don't. I'm not the one doing that. I'm just the one pointing out the hypocrisy of approach.
The entire argument trajectory is "this path is good, that path is bad, this one will naturally get better, that one surely won't get better enough", and the primary evidence given is deceptively inapt.
Yes, he'll fix the speedtest flag thing (which BTW, several people were accidentally using it wrong because the feature is unergonomic AF, but it's convenient to blame the user for "holding it wrong" right?). No, single-threaded encoding/decoding is a valid use case and not a testing methodology flaw.
> ... to be forthright about how significant codec optimization historically comes after adoption.
Well, this trend has now been broken, so it's irrelevant to mention it. Codecs are now expected to be optimized well before adoption. Take the case of AV2: the AOM folks are currently improving the reference encoder, libavm. SVT-AV2 has just been announced. This is the reality we're now live in. The fact we're having this conversation is an instance of that trend change.
> "I have no reason to believe they'd be better anyway".
Why are you now selectively quoting fragments? That sentence goes "...and I have no reason to believe they'd be better *than dir-pred* anyway." The part you omitted makes the entire difference. Implementing splines will improve JXL's efficiency (again, this was never contested), but it IS unlikely that they'll fully compensate for the lack of dir-pred.
You know why? I *did* take a shot at writing an automated splines implementation for libjxl, and I failed. Not because I didn't know what I had to do, but because I couldn't find a quick way to generate high-quality spline candidates that improve overall quality while making up for their size and encode compute overhead. I'm sure that there's a hypothetical clever way to do it, but the point is that the equivalent tool (dir-pred) is dead easy to implement in comparison, done countless times independently, and it just works.
Efficiently leveraging JXL's splines into the encoding loop is legitimately a very hard problem. No, problems of this kind shouldn't be this hard, and it's healthy to call this stuff out instead of pretending that, some time in the future, the "potential" of "alien technology" coding tools will somehow be untapped.
> You see, any unfilled area is specifically mentioned only to dismiss it as meaningless.
Oh, you're still doing the weird "subtext" thing... you know what? I'm done. Have a nice day.
> Codecs are now expected to be optimized well before adoption.
Maybe I hallucinated, but I seem to recall that your very fine AVIF work happened only well after major browser adoption. "Rules for thee, not for me", I guess. Talk about trying to pull up the ladder.
> SVT-AV2 has just been announced
Best of luck to it and AV2. That's a non-sequitur.
> The part you omitted makes the entire difference.
It really does not. Because...
> Implementing splines will improve JXL's efficiency (again, this was never contested)
your given position is that it will never be improved enough. That is the problem. That your position is that it will never be improved enough to be better than AVIF (I mean again, though it was previously until quite recently) is entirely the problem.
You grip tight to this idea now that if JXL ever improves even one iota then you're vindicated in having graciously allowed it to improve, but you have not granted it the capacity to be improved any bit more than AVIF.
> than dir-pred
This is what I'm talking about.
> I *did* take a shot at writing an automated splines implementation for libjxl, and I failed.
I'm sure you're very good and very smart, and you've clearly done very good things. You're not the only very good and very smart person, of course. But you do have very strong opinions about what is and is not worthwhile for the world to pursue.
> Oh, you're still doing the weird "subtext" thing.
What can I say. I'm used to very good and very smart people knowing about subtext, so I guess I assumed you would.
I don't know why you are being condescending to Julio. His work has pushed AV1 forward significantly. His opinions (specifically about dir-pred) are shared by others who are in the encoding space because they're objectively true. If a breakthrough for splines is invented, we'd be elated, but as it stands, dir-pred is in virtually every video/image coding standard since the late 90s for a VERY good reason. Dir-pred also has two decades of iterative improvement and research behind it. The argument is that JXL's lack of this well-established coding tool is a serious handicap that will make it very hard - if not impossible - to compete with AVIF within a reasonable amount of time.
AVIF was performing better than JXL/WebP/JPEG for a long while, even before Julio and Gianni's tune IQ work happened. AVIF successfully met the bar, tune IQ further increased that efficiency gap.
AV2 needs optimized encoders and decoders because it now needs to justify its spot among AV1 and JXL on the web. That's why avm, svt-av1 and dav2d are well underway.
Hi there! I'm Julio (co-developer of libaom and SVT-AV1's tune IQ). Here there are some points worth mentioning, because I'm catching a whiff of bad faith with your comment that honestly needs to be called out:
- The inclusion of his two proprietary encoders (Aperture and Iris) just serves to further support the argument that JXL encoder devs have work to do to perform at the frontier, while also proving you only need a person or two to do so. The two FOSS AV1 encoders in the compo (libaom and SVT-AV1) are enough to prove this. Given that blog posts often double up as a way to show-case projects, I think it's fair game to show off a bit. Also, keep in mind Gianni is just 21 and starting his career -- reporting such strong efficiency results across several image formats (AVIF, WebP, Aperture) is impressive and worthy of celebration by the community!
- Tiles in AV1 go through the deblocking filter, so there won't be any seams after decoding. In fact, AVIF encoding solutions (like libavif) enable tiling by default. If there were seams, people would've noticed those artifacts and yelled at the libavif maintainers.
- *Because* JXL doesn't have a deblocking filter, you could argue that JXL effectively decodes to numerous "mini-tiles" -- each one equaling the size of a coded block. And indeed, you WILL see those boundary artifacts when quality isn't high enough for EPF, Gaborish and/or LF smoothing to mitigate satisfactorily. This is what Gianni's post covers.
- AVIF scales very well under multithreaded decoding scenarios, thanks to the excellent work of the dav1d devs. The main conclusion wouldn't have changed -- AVIF is significantly faster to decode than JXL.
- In progressive decoding, a very valuable feature is "bytes to first usable image". That's what the comparison is focusing on -- it's not a cherry-picked point at all. By usable: you can tell the pass isn't a "blurhash", but you can actually discern each element in the picture with reasonable detail. You can play with the JXL demo yourself -- JXL roughly needs 3x as many bytes to get to where AVIF is in quality, and JXL is still a bit more blurry in general. This applies to every image in the demo, not just the poke bowl.
- The folks who coded the JXL demo happened to use two passes for progressive AVIF, but you can use up to four -- including adding an even lower-quality "blurhash" pass, and/or a medium quality pass. Yes, it's desirable to control the number of passes and quality at the encode stage.
It saddens me to see discourse on the internet about AV1's ancestry being focused on Google's VPx line of contributions, while diminishing those coming from Daala (entropy coder...) and Thor (QMs...), as well as inter-company collaborations (CDEF).
A redeeming outcome is that with AVIF's new image tuning modes (both in libaom and SVT-AV1), Gianni and I managed to utilize as many AV1 coding tools as possible, including QMs that sorely needed a well-deserved spotlight.
Also, thanks for paving the way to the current state of multimedia compression! I'm a longtime fan of Xiph(ophorus) since the 1.0 beta/RC Vorbis days. We used your image sets a lot during our testing.
You're welcome. That is what those image sets were for.
I was really happy Thomas Davies added quantization matrices to the spec, because if he hadn't, I would have had to. Fortunately, he did a good job. I am glad someone is finally using them. Don't sleep on per-segment quantizer deltas, either.
The quick answer is: you can have up to four passes. With progressive AVIF, you let the browser avoid rendering previous passes if a subsequent one has already been downloaded. It's more efficient and saves battery.
BTW, you can configure the AVIF encoder to have another in-between pass or two so the quality jump isn't as big. The JXL folks just happened to go with only two total passes.
Thanks to AV1's inter encoding, overhead is overall minimal as each pass can refine on previous ones. Think of it as a mini-video. Because of this, in the case of images with a lot of repeated patterns, progressive AVIF encoding can actually result in more efficient images!
Looking at the demo linked in the blog post, I can make a similar quality thumbnail for a smaller size than the delta between the static and progressive AVIFs
The quick answer is: you can have up to four passes. With progressive AVIF, you let the browser avoid rendering previous passes if a subsequent one has already been downloaded. It's more efficient and saves battery.
> A lower resolution image layered below the full resolution image, which is loaded and rendered first.
I'm curious, where did you learn progressive AVIF works like this? Have you actually read the spec, or does your understanding comes from somewhere/someone else and never challenged the truthfulness of it? Progressive AVIF is truly "progressive" -- it never involves "loading a thumbnail" or "layering an image over another".
In reality, each pass (up to 4) can refine previous ones (thanks to AV1's inter-encoding toolset), avoiding storing redundant information between passes. The viewing environment doesn't need to render a given pass if a subsequent one has already been downloaded. Finally, scaling is configurable -- you can have your first pass already be at full res, just at a lower quality.
Hope this helps clarify how progressive AVIF actually works under the hood.
Progressive AVIF is very flexible: it supports up to four passes, at configurable quality and dimension scaling levels. You can have any given pass reference up to two previous ones for refinement (thanks to AV1's strong inter-encoding capabilities), and you add filters to non-final passes (like blurring) to achieve a desired loading aesthetic.
That JXL page happens to use two passes, but the knobs are there to customize the experience to fit the use case.
I find it misleading to call AVIF's "up to four passes" "very flexible".
It seems quite limited compared to the JPEG XL ability to truncate the bitstream anywhere, or send the progressive updates for salient regions first [1].
> I find it misleading to call AVIF's "up to four passes" "very flexible".
Wow, what a way to misquote me. Let me repeat what I actually said:
> Progressive AVIF is very flexible: it supports up to four passes, at configurable quality and dimension scaling levels. You can have any given pass reference up to two previous ones for refinement (thanks to AV1's strong inter-encoding capabilities), and you add filters to non-final passes (like blurring) to achieve a desired loading aesthetic.
So, I re-iterate progressive AVIF is flexible, because:
- Intermediate passes in AVIF can look as sharp or blurry as desired -- you don't have that kind of control with JXL
- Intermediate passes in AVIF can semantically be different from the final pass -- very useful if you want to add a "loading" mark to the non-final passes to inform the user the image is still loading
- The four pass limit is A GOOD THING, as you want an image format to have a reasonable worst-case upper bound on energy consumption due to sum of partial decoding + display refresh updates -- there's such a thing as having "too many passes", and uncapping the limit would be irresponsible
- You can absolutely do saliency encoding in AVIF, as AV1's inter-frame encoding naturally allows for it efficiently
Ah, an accusation of misquoting. I actually quoted your exact words minus "is" and "it supports".
I think we disagree on the degree of flexibility, for sure.
A cap on layers (under user/browser) absolutely makes sense, but 4 at the format level is quite limiting, especially if you want to spend some of them on salient regions.
I agree that's possible, but not that it's efficient. You'd waste a few KiB on encoding skip blocks - AVIF layers represent the whole image, whereas JPEG XL can efficiently encode and update at group level.
How flexible did Jake find AVIF progressive in 2025? [1]
"it seems pretty limited. Only particular scaling values are allowed, and 1/8 is the smallest. Supposedly, additional layers are possible[..], but whenever I tried this, the encoder would error out, or explode the file size to ~400 kB, even at lowest quality. I guess that's why it's marked 'experimental'."
> Intermediate passes in AVIF can semantically be different from the final pass
I'm not sure if you're aware of this or not but juliobbv is the developer that fixed progressive AVIF, he is fully aware of how it works and how AV1 works in general.
"You'd waste a few KiB on encoding skip blocks"
This tells me you don't know how codecs work... skip blocks are not expensive to code, they are very cheap.
Jake's blog post is outdated by the way, the progressive functionality is much more advanced than it was at that point in time because Julio worked on it.
My conservative estimate is that a few-MP image with all but a smallish region encoded using skip blocks will spend a few KiB on that. This is very expensive compared to sending only a bounding box, hence it is unattractive for purposes of updating small regions with a whole-image layer.
I am glad to hear AVIF progressive has improved. But note that my original comment was: I find it misleading to call AVIF's "up to four passes" "very flexible". I believe that stands: contrasting the flexibility of (purpose-built) JPEG XL vs. the fairly strict limitations (inherited from video) of AVIF, I am astonished anyone would still call the latter "very flexible" by comparison.
Hi, I'm the author of the article you're quoting. AVIF was limited at the time of writing, but has since seen massive improvements, including progressive support.
Sounds like some strong assumptions here, particularly a stable and non-metered connection.
Imagine fast scroll across an image gallery on a slow connection (including cell handovers).
Or range requests, where a service worker only downloads the header+preview portion, and when clicking on the image, no need to re-download that.
Or even a browser that truncates all images, to protect users who might visit a page with huge background images that blows through their prepaid data plan.
JPEG XL anticipated, and accommodates, these use cases.
These seams are especially noticeable in the background, and align with coded block edges. You can tell the encoder is trying to conceal them as best as it can, but this cannot be properly mitigated without a proper deblocking filter.
For comparison, AV1 tiles normatively go through the deblocking filter.
reply