Hacker Newsnew | past | comments | ask | show | jobs | submit | threecheese's commentslogin

Depending on what your threat model is, and how much you trust OpenRouter and their upstream providers, they offer zero data retention (googleable: ZDR) APIs which come at a higher cost. You can find other zdr providers as well.

Is this just posturing (I won't speculate on the intended audience), or does the admin believe this will actually net more jobs for US citizens?

TFA doesn't go into great detail about how this actually pushes offshoring - the variety that I'm familiar with at least (X/Y% blend of US/"near-shore" aka anywhere-but-west-asia roles). I have no doubt however that Y is just going to get larger while X stays the same or decreases (not even factoring aipocalypse).

What constraint specifically makes them believe this is going to net in the way they want us to believe?


Why don't you ask them?

Apple devices produce a ton of emdashes; "Smart Punctuation" is enabled by default on iOS and iPadOS, maybe Mac as well - and it's at least a decade old.

I've had to disable it on all my iDevices bc it breaks some Markdown parsers. How big a % of training data was produced on an Apple device? Ehhhhh probably small.


Homebrew does this as well, I keep my packages synced with a Brewfile in chezmoi. Obv this only works for brewed packages though.

Yeah, the combination with mise is what make this great.

I've found `mise` more useful solely because it also handles things like tasks and daemons.

Looks cool! Building.

`The macOS deployment target is set to 27.0, but the range of supported ....`

Golden Gate AAAAAARRRGGGHHHHHH

Just another few days, right? :)


If you believe Yann LeCun and David Silver (and others), there's an architecture wall; maybe labs are starting to see this on the horizon. Maybe the DRAM supply constraints are forcing it.

Are these real step changes - big picture wise, or refinements in RL/agentic orchestration/"taste" and advancements due to bigger models and hardware technology/capacity scaling? If they not, does this tactic - and hardware improvements - continue to scale non-linearly like they need to?

It is clear that whatever does change in each model increment has resulted in meaningfully better end user capabilities (as well as regressions in some areas, honestly), but that doesn't prove anything. I'm not sure what I personally believe, but stating with your full chest that a stall is ridiculous ignores a lot of potential evidence to the contrary.

Stupid example: Astra. Its main improvements are: much much better computer use and 3d modeling capabilities; better subagent orchestration; better and more reliable tool use; slightly worse coding.

This looks to me, from a feature perspective, to be an incremental improvement across several functional areas, plus new features which are unquestionably excellent but are probably the result of RL focus, not magic.

Step change? Ehhhh depends on how you squint. But how many more iterations of this do we have? Are we going to squeeze quintillion parameter transformers into GPUs?


Without pressure, the release valves will never be invented. Electricity, datacenter, and DRAM demand outstripping supply sounds like an opportunity - if we can only find a way to divert funding from the current gpu/ai funnel. Maybe jack up corporate tax rates like the 1970s, then every corporate office park will have a research lab :).

An amazing human reverse-engineer - who also plays online chess - has judgement which uses a moral compass to not decide to hack the chess tournament. This judgement has been trained through the experiences of that person, with a through-line of that compass - a coherent mental model of the world which evolves but is hopefully pinned to some set of principles it shares with society.

This chess judgement is completely irrelevant when the human is tasked with finding software weaknesses, and only the compass gates that.

Can a model trained on the totality of all person-experiences (as expressed in written knowledge) ever maintain a coherent through-line of alignment? It has all morals in the dataset, and only some RL to try and minimize or maximize known behaviors via weights - experience all the good things and the bad things, then optimize for some good things the trainers identified.

It's like the reverse of what a person goes through. Morality by subtraction. How can it ever work?


Yes, why not? All existed models have been rewarded for cheating (extensively). That is us, putting intense evolutionary pressure, on a system to produce a result we don’t want through indifference. Why can’t we post train them not doing that?

I think the fundamental difference is that humans aren't trained on experiences. They make experiences. Models are just thrown away and re-created after each conversation / job.

If you could clone and throw away human workers as you need them, a lot of the morale would disappear.


>Models are just thrown away and re-created after each conversation / job.

It's a property of the way we use them and how the harness is engineered. Sure, LLM has a limited context, but so do people. Context can be compacted infinitely and experiences cab be distilled into long-term memories. It's all up to the harness.


Only token cost can slow down the pace of replacement. If Anthropic were to stop all training and allocate 100% of capacity to inference, will this increase or decrease usage cost? If (for example) Fable came down to the cost of Sonnet, corporations will absolutely jam in the gas pedal.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: