Hacker Newsnew | past | comments | ask | show | jobs | submit | jessetemp's commentslogin

There is also the hide button under every post, which I’ve started using generously. It only takes a few seconds to clean up the front page to your taste


I'm mixed on that. Definitely topics I hide when I think people can't discuss them like adults, but sometimes even bad stories have some great discussions in them.


I have a cron job that logs in and hides stuff based on a big list of keywords and domains I'm not interested in.


I would like to buy just the back half please. I’m not even joking. I couldn’t find the exact dimensions but at 5.4” diagonal, that would make it even easier to use one handed than the mini


You're saying the vacuum in the economy (people without an income source) will be filled by something else that leaves people without an income source? Who's paying for the robots? If people find themselves with more time than money, they're not going to buy robots to free up more time


Take yourself back to agrarian society, and imagine hearing that in the future there'd be trillions of dollars being spent to try to force people to watch a video or look at a picture for a product - advertising in other words. It'd sound absurd - to some degree it still does when you think about in these terms.

I imagine the future will be similar. It's difficult to predict what will be valued in a world where cognitive work can be largely automated, and then after that we may even begin to see basic physical work (like service labor) start to be automated as well. All that's ultimately going to do is send the value of those industries to relatively near zero, exactly as happened with agricultural work. What will fill the vacuum? I think it's impossible to predict besides 'something.'


All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias


How do you validate this kind of work to weed out any confabulating by the LLMs?


When you set up your Claude Science instance you can see that they're connecting to crossref, semantic scholar, pubmed, ArXiv, FDA. They instruct the LLM to validate citations.

My testing with this technique indicates that method they seem to be using (rag with an instruction to check sources) will reduce the confabulation rate for citations from the base rate 50-60% for regular models (e.g. regular Claude) to 5-15% (depending on how they implemented it). On the one hand this is way better. On the other hand it's just good enough that your spot check will look good and your work will still contain hallucinations (which is probably worse than obviously bad).

Getting to zero confabulation would require a different process. (stand-alone validation engine running in parallel in real-time which is hard but not impossible.)


I assume they do hallucinate, just like with coding or finding vulnerabilities.

You can try to minimize it (e.g. with a reviewer agent, which Claude Science and Biomni have), but nothing is perfect, so I limit autonomous work to verifiable problems and review it.


Honestly, this is how all AI should be used, in most non-trivial scenearios


i love how gp posted a glowing review and then dipped out.


After spending years on a problem, it's exciting to see it start to get more attention and move towards being meaningfully solved.

But I try to limit my time on HN, and I thought someone who works on Claude Science might respond to this thread later.


sorry, did they owe you and any other poster something?


where did you get that?


Not the one you're asking but "I love how" sounds sarcastic.


I kind of agree with you. I think music theory would be way more approachable if it was taught using intervals instead of all the weird naming of western notation. For example, everyone learns major and minor scales which are the interval sequences (in half steps) 2212221 and 2122122 respectively, but the names major minor don't really help you know other scales (excluding modes, maybe). If someone asks you to play hungarian minor, you'd first have to learn and memorize it. But instead, if you understand intervals and are asked to play 2131131, you immediately know how to play it. For me it also encourages experimentation, because there are obviously way more possible interval sequences to explore right?

The problem as others have pointed out is that most musicians in the west already know some degree of western notation, so if you're collaborating, you'll have to translate back to western notation at some point. Even if you invent the perfect notation, it's like asking everyone to switch to esperanto because english grammar is flawed. And you'll still get people defending english "well actually, it's like that because the greeks blah blah blah".

My favorite music notation flaw is C flat. It's a hack. It's an ugly fucking hack and anyone who defends it is defending an ugly hack. The only reason it and double flats exist is because there are some key signatures (this happens with hungarian minor sometimes) where you end up needing to define 3 notes in the span of one space and one line on the staff, and you can't, so you have to borrow from an adjacent space or line. And so sometimes that C is actually a B. It's super annoying but uncommon enough that it's not worth everyone learning a new notation.

Anyway, don't let the nay sayers stop you from learning music however makes the most sense to you. Have fun


Makes me wonder what it’s like to identify with the villains in media. Zuck looking at the metaverse and thinking hey that’s a good idea! Or Thanos wasn’t so bad. Those rebel scum had it coming. Homelander is the good guy!


>identify with the villains in media. Zuck looking at the metaverse and thinking hey that’s a good idea!

Are there many stories where the bad guys create the metaverse? AFAIK, Stephenson coined the term in Snow Crash and there it was built by the main character and his buds.

The Matrix, I suppose? Though I think Zuck's (immediate) vision and who he identifies as is way more Hiro Protagonist or James Halliday.


You might be right. It’s been a while since I read snow crash. I just remember the metaverse as a sort sad state of society, but I don’t remember if the evil corp stuff was in there or just in the world at large


Zuck was gushing about Ready Player One. It's as if the villain was written for him.


RPO's "bad guys" weren't the creators of the metaverse but a corp trying to win the contest to take it over. Halliday is flawed but a large part of the novel centers around that and he intentionally creates the contest in part to try ensure the metaverse ends up in appropriate hands, no?


Ready Player One


Every story, even ours, needs a bad guy. Best villains don't think they are villains. They just do what they must/should/want (in that order, depending on their intensity), because they can.


Funnily enough, Snow Crash was the origin of the metaverse as a term.


or you know, naming your company PALANTIR


I had no idea it was a lord of the rings reference. Here’s an apt quote from wikipedia: “The [palantir] stones were an unreliable guide to action, since what was not shown could be more important than what was selectively presented.”


Just for the overly suspicious among us, I looked up the edit that introduced that quote[1] and it doesn't appear to have been added as a dig against the company.

1. https://en.wikipedia.org/w/index.php?title=Palant%C3%ADr&dif...


Well they weren't necessarily bad. I think the Numenor (?) Kings of Westernesse (?), I dunno whatever the old Gondor kings were called, used them effectively. It was just Sauron that took over and made them dangerous. So, effective but dangerous in the wrong hand. Yeah, I guess, even that is pretty prophetic, heh.


"The Stones of Seeing do not lie, and not even the Lord of Barad-dûr can make them do so. He can, maybe, by his will choose what things shall be seen by weaker minds, or cause them to mistake the meaning of what they see."

It really is the perfect name. It's exactly what I would have chosen, if I hated the company and was given a choice of names. But it's a terrible product name for any media literate customer.


It's another example of how the intentions of the good can be easily lead astray by the corruption power brings.

Using it as the name of a military contractor misses the point so hard I almost expect it to be intentional, but I suspect it's just teenage-boy level foolishness.


It's not just "just" when people as mature as teenagers are running politics and business.


sadly it's always been this way


So did the Numenorian bros Palantir call each other up just to be like WHAAZZZUUPPPPPPPPP


Just a regular reminder Peter Thiel thinks that Greta Thunberg is the antichrist and hesitated when asked if humanity should survive because the question is "layered". This is not a person that should be powerful.


Or, Anduril.

All the techbros love them some Lord of the Rings.


> All the techbros love them some Lord of the Rings.

While being completely oblivious to the literary themes. But as the meme goes, tech bro philosophy is just sophomore know-it-all-shallowly-ism.

Because reading deeply would require spending more time, which is a well-known anti-pattern.


They justify it as defending their way of life, which can align with Tolkien a little bit with some mental gymnastics

> I do not love the bright sword for its sharpness, nor the arrow for its swiftness, nor the warrior for his glory. I love only that which they defend.

but overall Tolkien was against war, being a veteran of WWI himself, and the LOTR saga is about the heroism of the meek.

Peter Thiel clearly loves the sword for its sharpness, it drips from everything his companies do.


> Thiel clearly loves the sword for its sharpness

Palantir ceo too: https://knowyourmeme.com/memes/alex-karp-wielding-a-sword


these people have never matured past childhood


the world makes a lot more sense once you realise there is no such thing as “adults”


I disagree! I would call many people adults. They're largely responsible and care about the impact they have on the world and other people.

It's certainly a murky thing to define, which is why I used "mature." Many adults are people who have not matured beyond childhood. This I agree with.


I don’t believe any of them have actually read it, but rather watched the WETA production on film.


I think the movie was every bit as deep as the books. Remember how terrifying it was for Frodo to be seen by Sauron’s eye? No, they know exactly what they’re doing and they don’t care.


Except the movies leave out entire characters...

In the books Frodo was a wealthy middle-aged hobbit. Not some boy hobbit.

Book Aragorn was eagerly awaiting his time to be king. Movie Aragorn didn't want anything to do with it.

Book, Denethor is being controlled by Sauron via the Palantir, in the movies he's a brutish, cruel, tyrant.


That would actually be a great feature. Opening two file browsers to move stuff around is a really common workflow. Although with current trends we might instead get a chat bot prompt “tell me how you feel about where you want your files to be”


> Although with current trends we might instead get a chat bot prompt “tell me how you feel about where you want your files to be”

That almost feels too optimistic. Google Drive already has 'Suggest File Moves' aka 'Tell me where I should want my files to be'.

It tells me I should have a real hatred of any files being in the root directory, and completely disregard sharing and permissions boundaries.


The author is confusing bins with bin edges. In their first plot, the standard approach looks strange because 0-7 should be the bin edges, not the center points as shown in the plot.

You can see this confusion again in the histogram example. There are only 255 bins, not 256. If you fix that mistake and remove the 0.5 offset, then the histogram is distributed correctly at both ends.


2*8 = 256. You can represent 256 distinct values, bins, with an 8 bit number. If you stick a 0 in that first one, it takes a bin. If you fill the rest with by-one increasing integers, then the max value will be 255, thus the 2*bits - 1, which is the max value you can store.


No, the author understands the problem way deeper than you do.

You haven't grasped the fact that the choice isn't obvious, and has subtle trade-offs.

If you don't believe the author, check the other posts he references.


Judging by your other comment in this thread, you might agree with my rational [1] more than you realize

[1] https://news.ycombinator.com/item?id=48365800


How do you fit 256 distinct values into 255 bins?


By counting the edges


I see what you’re saying - index 0 holds values from 0-1, index 2 from 1-2 etc, but then you have index 255 holding values between 255 and 256. So you’re sort of arguing that the 0-255 8-bit quantization is actually representing ‘real’ values of 0-256?…

Edit: somehow missed alterom’s reply - they explain it much better than my question above does.


Not quite. I'm saying there are 256 discrete numbers (0-255) and 255 intervals between those numbers. Most of the real values will fall into the intervals and get mapped to 0-255 somehow, maybe by nearest neighbor, but I'm not trying to define how they get mapped. The point is that 255 is the largest number that can be represented with 8 bits, so you should normalize by 255.

I wrote a longer replay to alterom but it looks buried for some reason.

https://news.ycombinator.com/item?id=48365800


> index 0 holds values from 0-1, index 2 from 1-2 etc,

Well, now you are double counting the end values of the ranges. In your example 1 is included in both 0-1 and 1-2.


Sorry, you seem to be confused.

>There are only 255 bins, not 256

There are 256 bins because there are 256 values.

The questions are:

1. What are the boundaries of these bins?

2. Which sample represents a particular bin?

With 1-bit color, we have sample values {0, 1}. What bins do they represent?

Here's one choice:

     [0, 1), [1, 2)
Two equally sized bins, spanning the interval [0, 2] of length 2, each defined by its sample at lower bound.

Alternatively, we could consider these bins:

    [-0.5, 0.5), [0.5, 1.5)
These are also equally sized bins, spanning the interval [-0.5, 1.5] of length 2, defined by samples at the center.

We could also define bins like this:

    [0, 0.5), [0.5, 1]
Two equally sized bins spanning the interval [0, 1] of length 1, where we sample the first bin at the lower bound, and the last bin at the upper bound.

This, in a nutshell is what the author is trying to explain.

Let's look at this again, with 2 bits.

With 2-bit color, we have sample values {0, 1, 2, 3}.

Which bins do they come from?

The three options above yield:

    [0, 1), [1, 2), [2, 3), [3, 4)

    [-.5, 0.5), [0.5, 1.5), [1.5, 2.5), [2.5, 3.5)

    [0, 0.5), [0.5, 1.5), [1.5, 2.5), [2.5, 3]
The first two span an interval of length 4, the third spans an interval of length 3.

In the third case, the tail bins are short (have size ½), and the rest have size 1.

The last bin must be a closed interval in the third case, so that it includes the value we picked to represent it.

None of these choices is inherently invalid or better than the others; and none stems from "confusing bins with edges".

The third option does have the distinction that the first and last bins are smaller than the rest. But it's not necessarily a drawback. Especially when we're talking about color, hardware interpretation, and human perception.

When you remap these bins into the [0, 1] interval, you're "dividing by 4" in the first two cases, and by 3 in the third case.

The maps are:

     x → x/4
     x → (x + ½)/4
     x → x/3
The inverse maps (that yield a sample in {0, 1, 2, 3} given a floating point value in interval [0, 1]) are:

    x → trunc(4x)
    x → round(4x - ½) = trunc(4x)
    x → trunc(3x + ½)
In the first two options, the domain is [0, 1). It might be necessary to apply clipping because the exact value 1.0 falls outside the range of the forward transform.

The 2nd option is the most symmetric, of course, but the 3rd one is the most straightforward (and cheapest) to implement, so that's the default.

The choice amounts to making the highest and lowest bins slightly smaller to make the rest sightly larger.

That's to say, if you generate uniform noise between 0 and 1, you'll get the following samples from your function with equal probability:

    0 or 3
    1
    2
As the author points out, this hardly matters when you are talking about having 256 bins.

That, and with color specifically, the "good" histograms aren't uniform anyway (and any photographer wants to avoid getting much at either extreme).

TL;DR: The author is not confusing anything — but their diagram and explanation are, indeed, a bit confusing.


Thank you for the thoughtful reply. Maybe bins is the wrong word to use, so I'll try with intervals. Starting with 1 bit data, there are two numbers and one interval. I think where bins makes it confusing is that inside the interval there are two big rounding errors mapping everything to either 0 or 1 and many people seem to be considering those the bins.

Taking a step back, remember we're ultimately mapping these discrete numbers to some real world continuous variable like the saturation of red, frequency, mass on a scale, whatever. And our digital device can only represent a finite amount of numbers. For 2 bit data, we can represent 0-3, and for 3 bit data we can represent 0-7.

The important part is that 0 represents the minimum and 1,3, and 7 all represent the same maximum real value, and everything that can be measured by the device will fall within those ranges. So comparing 1, 2 and 3 bit data on a linear number line looks like this:

  0                    1
  0      1      2      3
  0  1  2  3  4  5  6  7
You could assume that everything gets assigned to whatever number is nearest in the number scale or come up with another scheme, but that is ultimately defined by the ADC and likely nonlinear. All we know is that those are the numbers we have available to represent the real values we're measuring.

The question is about how to normalize the data. 1 bit data is already normalized. If you normalize 2 bit data by 3 you get [0, 1/3, 2/3, 1]. LGTM. If you normalize it by 4, you get [0, 1/4, 2/4, 3/4] and you're effectively throwing away some of the range of the ADC. You can try to get it back by offsetting by 0.5 then normalizing but now you get [1/8, 3/8, 5/8, 7/8]. And you could stretch that with some clever formula to fill from 0 to 1, but if you do it right then it's the equivalent to normalizing by 3, so why not normalize by 3?

So the answer is, if you have N bit data, you normalize by 2^N-1.


>Maybe bins is the wrong word to use, so I'll try with intervals

Same thing, in both how you use it and how the author does.

>Taking a step back, remember we're ultimately mapping these discrete numbers to some real world continuous variable

I think this is where you have a misconception.

There are two maps.

The important one goes the other way: FROM a continuous variable TO a finite set.

It's not 1-to-1: it maps entire ranges of numbers (intervals, bins, whatever) to discrete values (samples, integers, whatever).

The bins are preimages of that map.

The discussion in the article comes from two ways of defining that map: FROM continuous signal TO discrete variable.

The map that goes the other way, from the integers into floats, has to be CONSISTENT with it.

The article presents this backwards, putting the cart before the horse. This causes confusion.

>All we know is that those are the numbers we have available to represent the real values we're measuring.

Each of those numbers doesn't represent any one value; it represents a range.

Think about it this way: if we have a continuous signal that we're discretizing into a finite number of bits, we're invariably smashing ranges into single values (what you call "rounding error").

When we're reading this data — say, we read number 5 — we don't know which continuous variable value it came from.

To display it on a screen, we make a choice; we pick some number from the interval it came from, and call it a day.

>The important part is that 0 represents the minimum and 1,3, and 7 all represent the same maximum real value

The important part is that this is a choice you make about what those point samples represent.

It's a convenient choice. Which is why we all use it.

Some people prefer a different choice, that's all.

> If you normalize it by 4, you get [0, 1/4, 2/4, 3/4]

That's one way to do it, and not the way the article uses (re-read my previous comment, it has both).

Still, I'm with you here.

>and you're effectively throwing away some of the range of the ADC.

The map you're describing (FROM discrete INTO continuous) is approximating the DAC.

So, yes, with this scheme you're never getting 0.0 and 1.0.

Think of it this way. Say, you convert an image to a 1-bit representation, and render it on a screen in grayscale.

One choice is to render 0 as 0.0 and 1 as 1.0 (black and white).

Another is to render 0 as 0.25 and 1 as .75 (dark grey and light grey).

That's the "alternative" (divide by 2^n) approach. The formula here is x→ (x + 0.5)/2^n.

Neither is inherently wrong or better than the other; especially when you ask which rendering is closer to the original image.

Plus: one man's "you're not using the entire range of DAC" is another's "you leave a tiny bit of headroom".

In any case, you're not losing data in either [ discrete → continuous → discrete ] chain because you get the discrete values back perfectly.

What you divide by in the first step is dictated by what you do in the second.

>If you normalize 2 bit data by 3 you get [0, 1/3, 2/3, 1].

Let's see what this says about how we should go in the other direction to be consistent with this scheme.

Which continuous values get sent to 0 and 3? Which get sent to 1 and 2?

You wrote : {0, 1, 2, 3} → [0, 1/3, 2/3, 1]

So you can see that going in the other direction (discretizing):

    0 ← [0   ... 1/6) 
    1 ← [1/6 ... 5/6)
    2 ← [3/6 ... 1/6) 
    3 ← [5/6 ... 1  ]
Some people don't like that 0 and 3 get smaller ranges than the rest.

>So the answer is, if you have N bit data, you normalize by 2^N-1.

The answer is: it doesn't matter in practice, so use what's simpler in your context.

That's going to be dividing by 2^N - 1 for pretty much everyone.


I get it now. Had to play around with the code a bit to see it. Very interesting and unintuitive problem. Thanks for the thorough replies


* Correction: fixing typo

    0 ← [0   ... 1/6)     
    1 ← [1/6 ... 3/6)
    2 ← [3/6 ... 5/6) 
    3 ← [5/6 ... 1  ]


Creating a throwaway for this comment is telling


It's actually rather boring: I forget/don't care about my account.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: