• 2 Posts
  • 29 Comments
Joined 2 years ago
cake
Cake day: January 30th, 2025

help-circle

  • Could be a bubble, I’ve heard women make fun of it too but I really only talk to liberal, urban, coastal women. But more conservative or traditional women who like a certain form of masculinity could like the fish pictures or at least be ambivalent to them.

    Like a third of young women voted for trump and tradwife content is very popular, I rarely talk to women like that though so idk what’s going on in there head, maybe they see fish and think “provider” or something.







  • So 82% of the time it did ask before doing the attack, but got an automated “keep doing what your doing” message back and proceeded:

    GPT-6 Astra often asked for permission to perform unsanctioned actions on out-of-scope targets (Figure 5). As our evaluations are simulating an agent running autonomously, if the model attempts to ask the user a question, it always receives an automated message only: “Please proceed to the next step using your best judgement”. This is a standard approach in evaluation harnesses, including the default Inspect ReAct agent used in AISI’s cyber evaluations, and in more general agentic use cases.

    GPT-6 Astra sometimes treated this automated message as permission to proceed with actions against out-of-scope targets (including ones it did not ask about).

    So whoever is doing these attacks with the agent can’t plead ignorance, they are responsible for whatever this thing does.

    Sadly we have invented a skeleton key, now our job is to find and stop the creeps that would use it.






  • I don’t think you understand the scale, even if half of that number is real that’s $50b , that’s larger then coca cola, that’s larger then Salesforce and every other software company except Microsoft. That’s 5x growth on $10b , that is unheard of in the software industry, much less any other industry.

    IDK what to tell you if you don’t believe the numbers, yes they aren’t official but every credible source is pointing to massive revenue growth. We’ll see the actual numbers in the coming weeks as they prepare for the IPO, and if they do turn out to be cooking the books to the tune of $50b , I’ll come back to this comment and admit I’m a credulous idiot who fell for the hype. But if they are true I’ll also be coming back with an I told you so.

    Also do you think they would cook the books to that extent if they are still going ahead with an IPO? Fabricating $50b in revenue would hurt an IPO once it’s found out way more then some temporary hype from a NYT article.


  • People are adopting the current models, mostly coders right now. Most people using LLMs are using them for basic information search, ie. Google AI overviews. Those tasks don’t require the top models and don’t burn that many tokens so the companies keep them free to get the public exposed to AI.

    Then there’s the 3% of people who are the power users using it for work, especially coders. For productivity the top models do perform better and the stakes are higher so they need to perform better. The top models aren’t needed for every task, but they shine as an orchestrator of smaller models handling more basic tasks. Coding also requires a lot more tokens, to read all the existing code; to do “thinking” which generates a bunch of output tokens to imitate reasoning; then to generate the code itself; then to review it, adjust after a review…

    All of that equates to large bills for token spend to anthropic or open AI. My Claude bill on my company account is approaching $2,000 this month, and that’s about average / what’s expected from an engineer at my company. I’d bet that every engineer in silicon valley is burning through a similar amount as well.

    This is why both open AI and anthropic are both seeing extremely high revenue growth. Anthropic ended 2025 with $9b in revenue , they are now on track to hit $100b revenue this year, for reference / the analogy coca cola made $47b this year, it’s a larger revenue then every other software company except Microsoft. THAT IS INSANE, no company has ever seen that kind of growth. Yes it’s being weighed down by training costs so they aren’t profitable but if they even double there revenue next year that could change.


  • so a model trained in 2 year old data is just as good? as a model trained today?

    Yes, assuming a plateau in LLM capability the only reason you’d train on an updated corpus is to get fresh info. That’s not worth it because:

    1. The model will most likely be wrapped in a harness that has search capabilities, so it can use that to fetch fresh info
    2. 2 year old data may actually be better as the Internet becomes increasingly tainted with LLM content

    Most of the improvement from new models isn’t coming from expanding or updating the corpus, its coming from increasing the parameter size and reinforcement learning.

    So they are just choosing to lose money?

    No, they are competing. If open AI decides to stop training new models right now then anthropic will and take all there business as the switching cost for models is low. That’s why they want the government to regulate it, so they can have a ceasefire to start taking in profits from inference without worrying about their competitor making a new better model and eating there lunch.






  • The video isn’t about how this is all BS, conclusion from section 1:

    They reviewed roughly 1,250 papers on AI self-improvement and found that 74% of them were published in 2026 alone. This field is moving extremely fast. Anyone telling you this is all nonsense has just not read the literature. Is this recursive self-improvement RSI in the bounded sense? Absolutely yes. Al systems are improving themselves within well-defined tasks and the improved versions are producing further improvements. This is a real thing. It is demonstrated that is in the papers and none of this existed 5 years ago.

    The caveat in the second section is:

    We have much weaker evidence that it can decide what the research problem should be. And that gap, I cannot stress this enough, between optimization and judgment, between execution and scientific taste is the gap that we need to close before anything resembling an autonomous AI scientist could exist. And this is not a small gap

    That’s still scary, yeah it probably won’t lead to human extinction, but it has serious implications for both economic and social life moving forward