The thread has little explanation as to what weird thing they’re doing to Codex that is making the default work poorly, and it kind of seems like it’s getting confused about whether it wants to set the caching mode or the breakpoint or both.
In any case, I find the behavior change interesting. It sounds to be like 5.5 and below may have been using a conventional attention scheme where a cached KV sequence can be easily used to restore a prefix of itself, but perhaps 5.6 is using linear attention or LSTM or another recurrent scheme where you cannot rewind the model state by just truncating it.
I have noticed the same thing starting with 5.6 when editing my last prompt inside the vscode codex plugin, I’ve seen the model’s thinking respond to the edit with a remark.
Slightly bummed out about it because in the past you could try different situations during a planning session and it wouldn’t pollute the cache but now it does. I’m not sure if forking the conversation has the same problem.
Then why are they (US frontier models) still so far ahead whenever I test them against the latest Chinese models? No bias here, I'd love them to be better for my own personal gain, but I haven't seen it
> It sounds to be like 5.5 and below may have been using a conventional attention scheme where a cached KV sequence can be easily used to restore a prefix of itself, but perhaps 5.6 is using linear attention or LSTM or another recurrent scheme where you cannot rewind the model state by just truncating it.
I feel like this is the kind of substantial change to your product that you would need to tell your customers about. It would be simply disrespectful to your customers to not disclose this upfront.
“Users are the product” is a phrase used when the users aren’t the ones paying for a free service. For a paid API the users absolutely are the customers.
My most awkward experience was a maintainer commenting on my feature request just to prompt a bot to "explain to issue reporter why this is very hard to implement."
It felt like they were trying to avoid me. They could have simply addressed me and given me the explanation they gave to the bot: it would have been simpler for him and more polite. I did in fact reply without waiting for the bot.
It's as though you're talking to someone and they were said to their 'assistant', "Explain this to this person" and walked away. It doesn't really matter what the explanation is, it's just gross.
The fact that companies with access to SOTA non-public AI keep having these kinds of dumb bugs in something as simple as a chat app helps validate my disbelief of people claiming AI is ready to replace all software development.
> The fact that companies with access to SOTA non-public AI keep having these kinds of dumb bugs in something as simple as a chat app helps validate my disbelief of people claiming AI is ready to replace all software development.
Except leadership doesn't care about dumb bugs. Never has, never will (until it's too late). It cares about velocity and cost.
They'd totally replace all software development with worse AI software development in a heartbeat.
> Except leadership doesn't care about dumb bugs. Never has, never will (until it's too late). It cares about velocity and cost.
This is the very moment at which I started my own business. I prefer to work for myself with some quality standards than to be in a rush in front of a prompt (not that I do not use AI at all, I do, but not for generating code most of the time).
I knew the future, at that time was basically: pressure for speed, taking ownership of course, even if they rush you. Wild-guess, probably with an AI, to add on top more trch debt. Make everything unmaintainable in the long term.
So this was the perfect moment to show that things can be done in another way and quality can be kept higher than the competition bc what I am seeing lately is people throwing things in a rush. Better twopieces of well-crafted software than 10 pieces of unmantainable junk.
Prompt edits leaking into the cache and affecting model responses is exactly the kind of billing-relevant behavior change that should be in release notes, not discovered by users.
Our codex on AWS Bedrock read / write cache ratio was less than 5%. Cache writes are very expensive and they were never being used. This results in codex on Bedrock causing ~10x what it should due to no caching and massive writes.
The workaround in issue resolved for me:
web_search = "disabled"
If you’ve got a workaround, I’d suggest updating the issue description to have it up top there so similarly impacted users can spot it quickly and benefit.
Codex usage feels exorbitantly high since today. They [0] are denying it, but the number of anecdotal users who decided to raise this as an issue (as a result it's trending on X) says otherwise.
My conspiratorial mind thinks they're doing this deliberately and using the resets to mask things so people can't tell their limits are reduced. The $200 / month plan covers about 2 days of usage for me right now.
Rookie mistake - it seems like they didn't follow manufacturers' guidance when installing the 10x engineers. One needs to clearly define which metric should be 10x'd before powering them up.
Applying Occam’s razor, which do you think is more likely:
1. OpenAI intentionally adds random overcharges.
2. OpenAI deprioritizes fixing actual bugs that cause occasional overcharges because doing so won’t affect their bottom line.
Here are the docs:
https://developers.openai.com/api/docs/guides/prompt-caching...
The thread has little explanation as to what weird thing they’re doing to Codex that is making the default work poorly, and it kind of seems like it’s getting confused about whether it wants to set the caching mode or the breakpoint or both.
In any case, I find the behavior change interesting. It sounds to be like 5.5 and below may have been using a conventional attention scheme where a cached KV sequence can be easily used to restore a prefix of itself, but perhaps 5.6 is using linear attention or LSTM or another recurrent scheme where you cannot rewind the model state by just truncating it.
Slightly bummed out about it because in the past you could try different situations during a planning session and it wouldn’t pollute the cache but now it does. I’m not sure if forking the conversation has the same problem.
I feel like this is the kind of substantial change to your product that you would need to tell your customers about. It would be simply disrespectful to your customers to not disclose this upfront.
Users is the product.
I mean i use AI too but was taken back when an agent popped up dictating what i should do and so on....felt weird
It felt like they were trying to avoid me. They could have simply addressed me and given me the explanation they gave to the bot: it would have been simpler for him and more polite. I did in fact reply without waiting for the bot.
Except leadership doesn't care about dumb bugs. Never has, never will (until it's too late). It cares about velocity and cost.
They'd totally replace all software development with worse AI software development in a heartbeat.
This is the very moment at which I started my own business. I prefer to work for myself with some quality standards than to be in a rush in front of a prompt (not that I do not use AI at all, I do, but not for generating code most of the time).
I knew the future, at that time was basically: pressure for speed, taking ownership of course, even if they rush you. Wild-guess, probably with an AI, to add on top more trch debt. Make everything unmaintainable in the long term.
So this was the perfect moment to show that things can be done in another way and quality can be kept higher than the competition bc what I am seeing lately is people throwing things in a rush. Better twopieces of well-crafted software than 10 pieces of unmantainable junk.
The companies you see struggling are ripe for disruption.
The workaround in issue resolved for me: web_search = "disabled"
If enough comments are added to the discussion, it might end up being collapsed.
[0] https://x.com/thsottiaux/status/2090675027670978569
My conspiratorial mind thinks they're doing this deliberately and using the resets to mask things so people can't tell their limits are reduced. The $200 / month plan covers about 2 days of usage for me right now.
"Random" accidents that always go against you, too biased to be random.
But don't notice that too much, you might start to see patterns here and there that you're not allowed to, might get you banned from places, etc.
1. OpenAI intentionally adds random overcharges. 2. OpenAI deprioritizes fixing actual bugs that cause occasional overcharges because doing so won’t affect their bottom line.