"If you’re searching for an air fryer, Amazon already knows quite a bit. They know the best-reviewed, least-returned, best-priced model. The only purpose of the ads is to get you to pick an air fryer that isn’t that one (or for the best air fryer, to keep you on track to buy the one you wanted in the first place)"
I don't get this part. Aren't those things unrelated? The aim of ads is to convince to buy regardless of the quality and other data points.
Yes. Technically, there are situations where the air fryer which is best for you and the air fryer that makes Amazon the most money are the same. However, with the number of airfryers on Amazon it seems safe to assume this is a rare coincidence.
What's the dataset used for this task? How does one prevent data leakage on the experiment itself? Are we asking about past events to predict the future?
I would use whatever you are comfortable with, I wanted a similar tool so I coded my own. Smaller API so that understand what is going on and it is easy not to get lost
Can't we just iteratively inspect the network traces then? We don't need to consume the whole 2mb of data, maybe just dump the network trace and use jq to get the fields to keep the context minimal. I haven't added this in https://news.ycombinator.com/item?id=47207790 , but I feel it would be a good addition. Then prompt it with instructions to gradually discover the necessary data.
But then I wonder, where the balance is between a bunch of small tool calls, vs one larger one.
I recall some recent discussion here on hn on big data analysis
Yes please, maybe there will be some solution that will fit the problem better! I recently released something similar, and because of the small API, I'm more comfortable using it.
I don't get this part. Aren't those things unrelated? The aim of ads is to convince to buy regardless of the quality and other data points.