Typesafe AI raises $870M at $7.5B
typesafe.ai
[24 comments hidden]
They may well be a good team to throw money behind if you are hoping to bet on a new AI lab.
[8 comments hidden]
[6 comments hidden]
[4 comments hidden]
https://benchmarkheaven.com/jev-models
According to this benchmark, Jev is currently trailing Quyet-1.0-Large and a few other hastily put-together LLM-based decision API-like setups.
[hidden]
[hidden]
I have no faith in the technique if it cannot do the basics (i.e. not real probabilities, the confidence for coin flip outcomes)
tried it a couple of days ago here: https://jevplayground.com
the "not real probability" disclaimer only appears after you get a result
[3 comments hidden]
If anyone else came up with the same concept on a Reddit thread (they have) it no-one would care without those characteristics even if you are "first".
Rebranding, execution, marketing, ex-<big_name_company> and mostly importantly, hype is what gets the investors scrambling into throwing money at you.
[2 comments hidden]
[5 comments hidden]
We've yet to see whether this is true, or is it just manufactured demand. There are dozens of Jev demos, but pretty much all of them are either cool but useless, or simply fake (i.e. harness doing 99% of the work).
[4 comments hidden]
its clear they own the mindshare around this type of primitive which is a massive premium
[3 comments hidden]
You can create a company with 2B shares and sell one share to your friend for $1000. Lo and behold, you own $2 trillion dollar company, leaving Elon behind.
[2 comments hidden]
Multiplying out revenue from the peak of their mini hype cycle while a dozen well financed competitors target them directly seems extremely optimistic.
[3 comments hidden]
Not them per se, but doomers.ai [0] which is a uh... "launch virality agency".
So less that Typesafe has a strong marketing muscle, and more that they paid at least $100K [1] for "organic" buzz.
[50 comments hidden]
[hidden]
When you are in the middle of a boom cycle, it's the hottest company that has the advantage. Investing in them is a matter of privilege and they get to pick and choose.
Also, Canadian VCs are bottom of the barrel as far as VCs go.
[3 comments hidden]
[2 comments hidden]
[hidden]
Acquired in the vc sense… not literal exit.
[2 comments hidden]
[1] not really but they did some innovative things and popularized a concept
[2 comments hidden]
[hidden]
[hidden]
Maybe in some cases. But counterexample, courtesy of The Information:
"It took just 15 minutes for Blue Owl executives to agree to invest up to $10 billion in future projects alongside real estate firm Primary Digital Infrastructure during their first in-person meeting two years ago, said Primary chief investment officer Bill Stein."
https://www.theinformation.com/articles/blue-owl-eyes-new-de...
AI seems to make some people lose their damned minds.
[hidden]
[19 comments hidden]
no, it was not
> was duplicated within a couple of days
was it already available or did it become available in a couple of days? it cant be both (neither is true, actually)
[18 comments hidden]
[17 comments hidden]
[hidden]
[15 comments hidden]
[14 comments hidden]
Jev is:
- accurate
- general purpose
- fast and cheap
Models we had before Jev had at most 2/3 of above qualities, but none of them were 3/3.
[7 comments hidden]
You're the one being deceptive here. Jev is trading accuracy, speed, and cost for generality. It's less accurate, slower and more expensive than trained classifiers. So it's still 2 out of 3, but with decimals. Maybe 2.2 out of 3 if I'm being charitable.
And the reason we didn't have that before is because nobody thought it's a good tradeoff.
[6 comments hidden]
Would Jev be more accurate in a specific task if it had been developed only for that task, as opposed to general purpose? Of course it would. So, sure, Jev is trading accuracy for generality. According to you "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier.
[5 comments hidden]
A business doesn't need Jev for the sake of Jev. Most business are solving specific problems.
And fine-tuning got a lot cheaper these days - I've seen claims here on HN that ~500 examples is enough to beat Jev.
> "nobody thought it's a good tradeoff", which again is false, there was huge demand for a cheap and accurate general purpose classifier
There wasn't. The hope is that there was a latent demand, but we've yet to see if it's truly latent or just manufactured.
Noone is saying "hell yeah, finally we got a general purpose classifier, my business needed it so much". The typical message is "this seems cool, let me see where I can apply it".
The fact that name itself is a play on Jevons Paradox illustrates that there was no demand until Jev was released.
[4 comments hidden]
[2 comments hidden]
Now, it could be viable if your business has literally hundreds of problems thats require classification. I just haven't seen those.
I treat the fact that almost noone was doing that as evidence that decision models aren't that useful/groundbreaking. That, and the fact that every single demo I saw was either fake (e.g. playing games), contrived, or plain wrong (e.g. using Jev for compaction).
[hidden]
The appeal and claim of LLMs for businesses is solving big problems.
LLM valuations to solve tiny problems seems iffy.
--
Are there significant problem domains LLMs are bad at that Jev is good at? Vs just 'Jev can do a subset of LLM things faster/cheaper'?
[2 comments hidden]
but you truly do sound like an angry 19 year old from your arguments.
- accurate - on what? on trust me bro benchmarks?
- zero-shot model are fundamentally general purpose.
- fast and cheap ; models on hf are FREE and fast enough.
[hidden]
[4 comments hidden]
> I was training custom ML models back in 2017.
Maybe you are a good person to ask my question then. I have not looked into Jev much, but is it much different from using a regular LLM and constraining its token output to the action space? (e.g. like using llama.cpp's GBNF grammars). Is it just that Jev's "confidence scores" are significantly better than the softmaxed logits? Or is there something else I am missing?[3 comments hidden]
I don't know how good Jev's "confidence scores" are, but I would be surprised if they were in any sense better than logits from some good LLM. One advantage of Jev here is that the confidence scores are easy to access. Most LLM API providers don't provide an easy/convenient way to access the logits. But that's a minor point, you could of course build something like this with LLMs (and many people have).
[hidden]
Also the situation isn't static. Investors know that the act of writing them a $870M check itself increases the chance that they'll be one of the winners, because that will attract more talent, customers, and funding to the company in a self-reinforcing cycle. And investors know that other investors know that, and that someone is going to write them that $870M check, so to some extent they're forced to think of the company as having already been successful at the fundraising and already having that momentum boost.
Only a small number of investors in the world can play the game at this level, because you have to smart enough to be right (often enough), and you have to be established enough to see the deals (be on every CEO's short list - because CEOs are only going to seriously pitch 5-10 VCs on a hot deal, if that). Otherwise you can't pull it off. Martin Casado and his team are among the few that can and I think their results reflect that.
[6 comments hidden]
Anyone know when this "have no moat" meme appeared? Even 5 years ago I don't remember seeing it on every post.
[hidden]
I would argue Dropbox did have a moat. It didn't merely store your data. It made it possible to make backup efficiently when bandwidth wasn't all that good.
Reading "no moat" so often is also tied to the fact those companies happen to be getting surreal valuations, at a quite early stage, showing no profit, building a tech that doesn't seem difficult to reproduce.
[hidden]
[2 comments hidden]
Instagram: I have to move all my posts and also convince all my network to move over
Github: not a terrible example
If I decided to move off of jev tomorrow it would be an api key and a base api path update. Maybe 30 seconds of work.
[hidden]
You can say the same thing about OpenAI/Anthropic, just an endpoint update, maybe a harness change if you use the CLI.
People have in fact been saying "what is the moat of OpenAI/Anthropic", yet here we are, trillion dollars valuation.
[hidden]
And if you could put the words "Bay Area" or "Stanford" or "San Francisco" next to your name... different story.
The VCs are not buying the idea or the tech, they're investing in the people. And they invest in a formula that has already worked for them before to make big coin. Prop somebody up, let them hire like crazy, and then get them get acquired, and then cash out. They don't care if it fails if they can make it succeed 1/200 times.
Canadian investors want you to have already succeeded before they help you succeed a tiny bit more.
[3 comments hidden]
There's your problem. The single biggest thing every Canadian VC is trying to figure out is "why are these people asking us for money when if they were any good they'd be in the US" so by simply asking them you're already signalling something bad. A lot of their enthusiasm for process is based on this suspicion and also that the entire industry is just a way for various professional services to extract most of the investment money, since that's the game they're so used to playing with the government.
There are some Canadian VCs earnestly trying to improve but they are overwhelmingly hilariously conservative and focused on unimportant signals over reality. This is one (but not all) of the major factors that drive basically every remotely ambitious Canadian company to run a corp in Delaware and go for funding from the US. The tax situation is the other major contributor.
[hidden]
[hidden]
But that does nothing to make up for the terrible investment community. Getting started here requires already being started.
When I briefly worked for a Toronto startup, it was like all of them went to the same private boy's schools together as kids. It was a status club.
I jumped ship to an American startup and made almost double the money dealt with 0% of the bullshit and they were bought by Google the next year.
[2 comments hidden]
I think the key differentiator was that a team found a whitespace in what ChatGPT was doing, main comes from the same pedigree and team is as conscious of marketing as their product. SF VCs love these out of the box challengers, and people are claiming to replicate doesn't seem to matter.
The amount raised feels surprising but again entire SF/US AI scene is primarily "add moar layers and GPU" one trick ponies at this point.
[hidden]
"Throw a bunch of money at copycats in the trend of the day" is very normal in SF/SV for the last several decades, going back to the dotcom boom.
The definition of "a bunch" has changed but so many "crypto for X" or "recommendations for X" or "uber for X" or "social for X", etc, things raised amounts that seemed wildly divorced from their market position.
[hidden]
Yes, anyone can wrap a decisions API around an LLM, but so what? If you want to compete then you need to compete on price, and it's not clear if OpenAI and/or Anthropic are able or willing to do that without building a custom architecture, and even then is a race to the bottom on pricing really what they want to pursue?
I'm not sure if OpenAI have announced pricing for their Decisions API, but they have said it's based on Luna which costs $0.10/M input, not even remotely competitive with Jev's $0.04/M input, which I'd expect has some headroom built into it.
Assuming that the architecture behind Jev is not just an LLM, and gives them some inherent efficiency/cost and speed advantage, then the question is whether OpenAI and Anthropic really want to duplicate this and have a race to the bottom on pricing for what may be a large part of the business automation market they are addressing. Is that what they want as their IPO pitch - we're selling potatoes, and think can grow them cheaper than Typesafe ?
[hidden]
There needs to be cleansing with fire. Weeds need to die. Trees need their branches cut. The sooner the better.
A shitty ass random "ai" startup built entirely on hype and astroturfing should not be raising anywhere near this amount of money.
[9 comments hidden]
[hidden]
[hidden]
https://www.youtube.com/watch?v=xNgQtzEl4lY
Jev is used as an example of a successful marketing launch where they worked with many X "creators" prior to its release, so that all the creators would repost to put it to the top of everyone's feed. Then, over the following days they'd repost so it maintained momentum.
See: doomers.ai, clickstrike, growth matrix, etc. They use coordinated engagement, paid influencer networks, customized messaging, etc.
Jev isn't a terrible product, but it's way overhyped.
[2 comments hidden]
[hidden]
If I seriously needed a classifier, I would just train my own and it would run on an iPhone. Any CTO worth their salt would suggest the same, because it's not even remotely comparable to training a large language model (w.r.t. compute or training data required).
[3 comments hidden]
Obscene marketing is all you need it seems.
[2 comments hidden]
Kind of sad seeing so many “tech people” (including here on HN) falling for it hook, line, and sinker.
[hidden]
[2 comments hidden]
[hidden]
I am integrating it into the product I am building and to me it doesn't seem like there is much need to go with a SaaS for this since the requirements are so light. I just can run it in Cloud Run and get all of the scale I'll ever need, and I get to tell my customers their data never leaves my environment.
[hidden]
[0] https://jobs.ashbyhq.com/typesafe-ai/9a94651c-5d63-4e82-8854...
[hidden]
[4 comments hidden]
I see a lot of people parroting the quick open source alternatives as being better on the benchmarks, but it's such a new category that I'm not convinced we have solid benchmarks.
I'm hoping a company releases an internal eval benchmark for these options. I'm sure some of the open source ones are solid in some cases, but would love to see more reliable data.
[3 comments hidden]
[hidden]
[hidden]
OK, I'm jealous... Lol
[hidden]
Open source[1]. It comes with 68 pre trained classifiers which run and train on CPU alone. They run locally, are faster than Jev/Laya/Decisions, and they can perform many different tasks; from labeling email, all the way to playing Doom[2]
[0] https://jeffyclassify.com/
[9 comments hidden]
[6 comments hidden]
[5 comments hidden]
[4 comments hidden]
VC is a hits business. Just one hit pays for 9 that didn't work out.
[2 comments hidden]
[hidden]
The difference in hit rate matters enormously and the existence of a hit tells you nothing about the denominator.
The size of the bet matters too - you could wipe out a hit with one bad pick if that bad pick is big enough!
If you don't have an estimate of the rate that you trust, you're just throwing money at dreams.
But it's WAY harder to get rich by being a pessimist than by being an optimist. And if you're a VC choosing where to invest primarily-other people's money, then you have no particular reason to try to talk them out of the hype.
[hidden]
[3 comments hidden]
[hidden]
Invariably near-AGI systems created by OpenAI/Anthropic will be very destabalizing. In the end the world will probably regulate AI capable of [any] <-> [any] input/output types. Models will need to be limited on their outputs by law so they cannot have unbounded, unpredictable outcomes. Jev is the ideal version of "benefits of AI without making humans obsolete" that might be the consensus once the track superhuman AI and its consequences are clear.
[hidden]
[hidden]
[6 comments hidden]
[3 comments hidden]
[hidden]
I actually resized my browser thinking maybe something weird was going on with flex-wrap or overflow or whatever it is.
[hidden]
For what it’s worth, however it works out, my guess is that the primitive Jev provides is likely to be considered essential in the future development of software.
[2 comments hidden]
[hidden]
[hidden]
[2 comments hidden]
[hidden]
If being an “AI Researcher” is a ticket to multimillion dollar salary, AI training talent cannot be contained to a handful of companies. It’ll become more common and diffuse. The old advice of not fine tuning, because it’s hard, goes out the window as that knowledge diffuses through the industry.
A similar thing is happening in search. For a long time labs have trained tailored embedding models. And now companies like SID training their own agentic models that are smaller and faster at search than GPT-5.
[4 comments hidden]
[2 comments hidden]
A lot of people are shouting about how Jev hasn't actually differentiated itself, but I question how much folks are actually experimenting with what's out there before coming up with an opinion.
For us, it's cleae that OpenAI rushed this out to meet the hype in the market right now without having a product that actually meets the bar Jev has set.
[hidden]
I did. Originally I had a project that I had been wanting to do and thought to use a decision model for it. Jev, OpenAI, etc. are all within percentage points of each other.
Then I used traditional ML and found a small classifier (gemma 4) with traditional embeddings worked 2x as well.
Jev is the general purpose ML pipeline for when you want average results. Nearly every application has a "better" option available with a small amount of work.
[hidden]
[hidden]
[2 comments hidden]
Unless china takes leadership in frontier space the picture is next :
1. cheap workhorses for classification, routing, other scenarios : Jev 2. coding agents with less erros : Anthropic/Openai, etc. 3. Science /Legal/Medical : A mixture of Jev+Anthropic scenarios
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
Curious to see how all of these can work together, and giving LLMs the ability to deploy their own workflows and built type 1 systems as they go has also been fun and useful.
[hidden]
[hidden]
[hidden]
[hidden]
[hidden]
You won the competition with VCs
[hidden]
[4 comments hidden]
We released an Apache-2.0, open-weight 4B decision model that scores above Jev 1.13 on JevBench's composite score (72.5 vs 71.5) and is currently the top open model there: https://benchmarkheaven.com/jev-models . Newer models coming even larger than beat Jev in intelligence as well.
- Same contract as Jev: state + typed questions in, calibrated probabilities out, one forward pass, no generated tokens. - Your data never leaves your environment, and there's no per-call fee.
Weights, card and run instructions: https://huggingface.co/h2oai/h2o-lightning-4b
[hidden]
One note on their speed comparison: the JevBench board shows self-hosted models with an adjusted latency of "2x + 0.15 s (assumption, not measured)". Our measured p50 is 29 ms; the adjusted figure is 0.21 s. They report 85 ms for theirs.
armcat[125 comments hidden]
EDIT: As I wrote this Microsoft just released their own Decision-1 model [3].
[1] https://developers.openai.com/api/docs/guides/decisions
[2] https://unsloth.ai/docs/basics/train-your-own-decision-model...
[3] https://commandline.microsoft.com/microsoft-decision-1-model...
dkersten[24 comments hidden]
That’s a different set of properties in terms of size, cost, and latency. Jev (apparently, not like I’ve seen its insides) is extremely cheap, extremely fast, doesn’t cost any output tokens as it speaks the output natively, can’t get the output wrong because it speaks the format natively, and (presumably based on the docs), the context is separate from the question, meaning it should be immune (or at least highly resistant) to prompt injection attacks.
It’s not just about the accuracy of the result, it’s a collection of all the properties that make Jev interesting.
Jev took years to develop, I strongly doubt that a copycat that was put together within days after Jev’s release will be able to match it on a sun of its properties. Even if fine tuned LLMs can outperform it on raw accuracy.
devin[11 comments hidden]
baobabKoodaa[10 comments hidden]
homarp[9 comments hidden]
and https://arxiv.org/abs/2510.01237
baobabKoodaa[7 comments hidden]
"SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization"
Jev is a general-purpose thing. That is a specific-purpose thing. General-purpose thing is not the same as specific-purpose thing. What makes people think these are the same thing? I don't get it.
devin[6 comments hidden]
baobabKoodaa[5 comments hidden]
Copypasta:
• A reinforcement learning architecture specifically designed for sales conversation analysis and conversion prediction
• A synthetic data generation pipeline leveraging GPT-4O to create diverse and realistic sales conversations
• Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features
• A meta-learning approach enabling the system to express confidence in its predictions based on conversation similarity to training data
• Integration mechanisms providing real-time guidance within existing sales platforms
• Extensive comparative evaluation demonstrating significant performance improvements over LLM-based approaches
rpdillon[4 comments hidden]
Jev is basically the embeddings side of an LLM. Yes, it's a good idea, but the moat is non-existent.
baobabKoodaa[3 comments hidden]
devin[2 comments hidden]
baobabKoodaa[hidden]
> Jev is basically the embeddings side of an LLM
This is how the paper you are referencing is describing the "key contribution" as it relates to embeddings:
> Novel state representation techniques using Azure OpenAI embeddings (3072 dimensions) with sales-specific features
And now you are claiming that the author of the paper created:
> "generalized" version of same
It's a bit hard to guess what you are trying to say, but if I were to steelman your argument, I would guess that you mean: training a model with a vocabulary that has some tokens representing things like "OPTION_A", "OPTION_B" is in your mind the same thing as "generalized version of sales-specific features in embeddings"? Is this what you were trying to say?
zwaps[hidden]
Literally misses the point of Jev, which you don't need to fine-tune to get accuracy nor - and no other model has this - some sort of out of sample calibration
hbrn[6 comments hidden]
But why? If the simplest way to achieve Jev's capabilities (accuracy, cost, latency) is by fine tuning a small model, what makes you think that this isn't exactly what Typesafe did?
And even if they did something different - what makes you think it was a good idea in the first place, given how easy their results were replicated without any "secret sauce"?
dkersten[5 comments hidden]
hbrn[hidden]
Plenty has been said about this claim. If you're still falling for this, I feel sorry for you.
If you remove wheels from your car, your car will get a "no speeding ticket" property, and yet there's nothing exciting about it.
> You can’t get that with a fine tuned LLM
Of course you can. All these claims are nothing but marketing.
user43928[2 comments hidden]
The only difference I am aware of is that probabilities are better calibrated with these decision models compared to regular LLMs which can output hallucinated numbers where your schema allows a number.
dkersten[hidden]
LLMs can produce structured output by limiting the next token based on a grammar, that ensures correctly formed output.
But (presumably, again I haven’t seen Jev’s insides) Jev doesn’t need to do that, it just has to output probabilities for each answer, and can do it natively without artificially limiting output tokens.
So jev is cheap, fast, never produces malformed output, its output always carries correct semantic meaning (but this doesn’t mean it always answers correctly), and it separates correct from prompt (which if it truly does that is the most exciting part).
PunchyHamster[hidden]
TeMPOraL[4 comments hidden]
mrinterweb[3 comments hidden]
lifeisloving[hidden]
Also Jevs purpose isnt to become its own thing. It will get aquired in 18 months by one of Andressen Horowitz's incestuous circle of companies and everyone will make money, and the person who buys it wont necessarily care if Jev itself makes them a ton of money. They're just passing chips around the table.
TeMPOraL[hidden]
creatonez[2 comments hidden]
dkersten[hidden]
The jev documentation says this is what you should do. It doesn’t mean that it’s actually designed for prompt injection resistance, but it could be. We won’t know for sure until they either release more details or someone proves otherwise.
But the split is one that doesn’t exist in normal LLMs and it’s the exact split that is needed for immunity or resistance to prompt injection.
mlmonkey[2 comments hidden]
To run the SDK examples below, use these OpenAI SDK versions or later: Python 3.26.0,
I thought Pythin 3.15.0 just came out, 3.26.0 must be really far off?
bayesianbot[hidden]
scottyah[2 comments hidden]
Have you heard that from a different source than OpenAI? From what I'd heard other models haven't gotten close, and the open source ones are like running gemma4 E2B against Opus 5.5- sure, the API calls go in and are returned the same but the quality isn't close.
rockwotj[hidden]
nlpnerd[hidden]
baobabKoodaa[4 comments hidden]
no, you can't, and it's unclear why you would think this.
ricericerice[3 comments hidden]
*if you have a sufficiently sized and quality dataset for the specific classifications you're targeting
baobabKoodaa[hidden]
9dev[hidden]
Then yes, easy!
ttul[19 comments hidden]
Oras[18 comments hidden]
I give it to them for creating the hype (good marketing), and for making a useful classifier. Not sure what they would scale rapidly though.
brink[hidden]
rsalus[2 comments hidden]
SaltyBackendGuy[hidden]
ModernMech[hidden]
mococa[2 comments hidden]
Actually they stolen the idea from a paper.
shdh[hidden]
tiborsaas[5 comments hidden]
throwaway7783[2 comments hidden]
PunchyHamster[hidden]
tyre[2 comments hidden]
Jev has been around for a couple weeks. Cost and performance matter more than anything. Staying with an existing provider (the # of people choosing Jev without already having a frontier API key is probably zero?) is way easier than this.
What the hell are we even talking about.
wordpad[hidden]
clickety_clack[4 comments hidden]
jeremyjh[hidden]
real0mar[2 comments hidden]
Espressosaurus[hidden]
OtherShrezzing[2 comments hidden]
People who can find things of value, and execute on those things, are worth investing in - even if that thing of value is replicated in short order.
Nobody was looking at decision models, but now that Jev has surfaced, everybody is.
vasco[hidden]
Investing isn't binary. You can be investible but not at a 7.5b valuation.
doctorpangloss[hidden]
By all means, become an A16Z LP.
amelius[hidden]
827a[33 comments hidden]
There's the potential for an inverse LLM play. In contrast with LLMs, all that seems to matter is the model, and the products the labs build around the models are all really samey and boring; the same left panel list of agents, main view agent conversation, right hand extra context, and we're now in the era of everyone creating the same cutesey furry friend on top of all this tech.
nlpnerd[3 comments hidden]
https://seldon-ai.com/blog/fronter-llms-are-semantic-interpr...
TeMPOraL[2 comments hidden]
nlpnerd[hidden]
Constraints can be an advantage in many systems. Consider typed and untyped programming languages. Many advantages with typed languages over the latter in terms of development and efficiency.
porridgeraisin[3 comments hidden]
Precisely this. Should be top comment.
Also, the model moat is understated as training data for these purposes also accrues to the winner, which due to the first mover advantage as well as the distribution advantage you speak of, is typesafe. In contrast to relatively open coding data. Openai anthropic also have that, but like you say its a different business.
outofpaper[2 comments hidden]
cheesecakegood[hidden]
If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.
pantelisk[9 comments hidden]
But even if others surpass them and make better solutions, the fact that nobody was able to see it before typesafe is a testament of what they might be able to come up with next.
I sound like a fanboy but I swear I 'm unaffiliated with typesafe. I was building my own version of this way before they announced JEV (mine was ALE and it was mentioned here on HN for a bit), in use for VR gaming (so one can give commands to NPCs with voice and supports multiple commands in sequence in a single pass), but I missed the "killer usecase" of being a new primitive, like everyone else.
TLDR: There's a lot of value in thinking ahead and seeing the future. The clones are nice and exciting but they give me "I could have built this first, yes but you didn't" vibes. I hope they manage to keep it up and push the space forward again
aranelsurion[hidden]
Here if anyone is interested: https://pantel.is/projects/ai-gaming-companion/
overfeed[hidden]
Counterpoint: there are a lot of one-hit wonders, and they vastly outnumber the idea-factory people. This is not to minimize those people, a single idea can be very successful (see Zuckerberg), but it doesn't mean your subsequent ideas will also be great (see Zuckerberg)
skeeter2020[hidden]
This all sounds intelligent and likely, and yet we can come up with countless counter examples where the first mover is not the big winner, and nobody cares about who did it first. There is typically way more value in nailing the execution of a big idea someone else came up with, rather than "seeing the future".
user43928[5 comments hidden]
Yes, this generates probabilities over a set of given options instead of the whole token vocabulary.
I don't see what's so exciting about it, compared to regular LLMs which support structured output options in the API.
senordevnyc[3 comments hidden]
I keep evaluating Jev for my product because of the hype, but the reality is that for my tasks, Luna is more accurate, only a little more expensive, and the latency doesn’t matter. I’d rather spend the extra $50 / month or whatever than have to shoehorn in another API and provider, and also lose the ability to change reasoning level and get reasoning summaries for eval purposes.
jeena[2 comments hidden]
senordevnyc[hidden]
But that’s just for my product.
pantelisk[hidden]
Models like JEV fall in the golden middle. Somewhat cheap, somewhat fast, somewhat general.
Many people have replied saying that narrow classifiers (eg bert) are better, and they are better in many ways (essentially free and faster). But... the real world is messy, in my experience it's actually very hard to make a really good classifier that will work across a specific niche domain especially when there is very fuzzy input. And the most interesting real world scenarios are nuanced and overtime there will be edge cases discovered where it fails and that will require retraining and re-evaluation etc etc. It can becomes a full project on its own that eventually collapses into a game of whackamole (improved in some direction but regressed elsewhere). Things like JEV (or similar models) provide ease of mind, just off load the complexity to it and move on to the next challenge type of thing.
nico[4 comments hidden]
This is a very important insight. And it applies to LLMs as well. Very few people were impressed with the capabilities of GPT 3, it was mostly a techie novelty
But then when they added chat on top of gpt 3.5, all of a sudden it was a huge hit. Sure there were improvements in the model from 3 to 3.5, but the biggest impact was from the chat experience
Conversely, when they created Eliza, a basic chatbot more than 50 years ago, people even got addicted to it, despite it’s ai model being something super rudimentary and basic compared to what we have now. The model capabilities didn’t matter as much as the experience the chat created
vrighter[hidden]
Darmani[2 comments hidden]
dannyw[hidden]
“The following is a conversation between a human and a helpful AI assistant.
Human: [prompt]\n AI:”
You would get consistent conversational chat, with the main limit being the model’s limited context window, and obviously frontier intelligence at the time.
I played around quite a bit with davinci-003 and conversational systems before ChatGPT. I kept using this format and API for a while, but at least during the ‘free research preview’ era, found it generally smarter than chatgpt. 3.5-turbo was probably the watershed moment where I moved away from completions on base models; to chat completions.
If you’d like to emulate this experience, spin up a base (not instruction tuned) checkpoint of Llama 1; or a more recent base model. You might underestimate how much ‘intelligence’ you can get :)
Back then the API exposed a lot of controls from sampling to logits, so you could specify a custom stop token (some rarely used Unicode); then switch to near-greedy sampling for more reliable “function calls”, etc; and then switch back to sampling for chat.
In the early days, before releasing something with the API, you had to get your use case approved by OpenAI, like the App Store review process.
olalonde[12 comments hidden]
Why?
porridgeraisin[5 comments hidden]
zhivota[4 comments hidden]
majormajor[3 comments hidden]
I haven't personally yet found a net-new-capability bigger than that from LLMs here.
sks_15[hidden]
porridgeraisin[hidden]
These guys also love text to sql. I know 4 people each implementing text to sql for their separate companies. Or text-to-redash dashboard in some cases. Funnily enough, text to sql is one of the things where LLMs are clearly sub-human. Mostly due to lack of easy verifiability.
reexpressionist[2 comments hidden]
For decision-making with neural networks, we instead need the older idea of estimators of the predictive uncertainty with constraints in the feature-representation space (over training/support of the estimator), as with Similarity-Distance-Magnitude estimators: https://pypi.org/project/reexpress-sdm/
SwellJoe[hidden]
rebyn[4 comments hidden]
SalariedSlave[hidden]
the claim is extraordinary, yet the answers cover only the mundane.
iinnPP[hidden]
The sum of people not even close to the topic has risen, according to my feels at least.
Has this rang true for anyone else?
PunchyHamster[hidden]
baxtr[hidden]
What I haven’t seen answered is why it isn’t reproducible within very short amount of time and low effort.
Sure, the large labs are the lazy incumbents at this stage. But any other startups would be able to copy the experience. What’s the barrier of entry that I am missing?
moralestapia[hidden]
hkalbasi[hidden]
Jev is 42$/B but OpenAI is 100$/B token.
bushbaba[hidden]
such an arrangement can end up beneficial to the VC firm
throwaw12[hidden]
But, for OpenAI this is not a primary business, for open source models as well, so they will not be chasing the market and customers to buy their product and promise them to maintain it.
TypeSafe will do all this, they will try to understand your use cases and then solve your pain point, while others are providing raw material.
jgilias[10 comments hidden]
phalangion[hidden]
hbrn[8 comments hidden]
Here's a hint: confidence is not generated by a model.
jgilias[5 comments hidden]
Do you think though that Luna being a model post-trained for chat produces over-confidence in logprobs?
hbrn[4 comments hidden]
Typesafe claims that Jev is calibrated, but there are plenty of examples where it completely fails (predicting die roll being the most obvious one).
Unfortunately calibration is hard to benchmark.
shados[3 comments hidden]
If you instead give it a list of probability for each number and ask it whats the probability of each number, the result will be accurate.
hbrn[2 comments hidden]
Did i hear that correctly? In order for Jev to be accurate you have to give it the answer before asking for the answer?
(btw this is exactly how Jev is playing games).
jgilias[hidden]
What am I missing?
As in, neither LLMs nor Jev are truth engines. Truth comes from the provided context.. Plus weights.. kind of fuzzy, sorry I’m thinking out “loud”
adrian17[2 comments hidden]
hbrn[hidden]
But confidence value is just a function applied to probabilities. It is not coming from the model, and it carries no additional information.
It is documented btw, and yet you will see plenty of claims that Jev is better than LLM because it returns both.
girvo[hidden]
user3939382[hidden]
nico[hidden]
* https://jeffyclassify.com/
* https://playground.jeffyclassify.com/#doom
* https://github.com/nicobrenner/jeffy
rpdillon[hidden]
shados[2 comments hidden]
Guess the (investment) market has spoken.
nacs[hidden]
The Microsoft article GP linked even shows the MS model having 95ms latency.
Even on price Jev being matched (the same MS model is "Input tokens cost $0.042 USD per million tokens. Output tokens are free.", same as Jev).
Cloudflare's Clef-flash model is actually slightly cheaper: "$0.038 in / $0 out per 1M" too.
Jev is being matched or exceeded in performance and price within a month of them going public.
shmoogy[5 comments hidden]
lifeisloving[3 comments hidden]
Fortunately Jev is cheap, so I dont think it matters too much, but I think its robbing people of the opportunity to learn and implement this themselves.
Also, I dont really want 3 companies responsible for censorship/classification.
jofzar[hidden]
nlpnerd[hidden]
nwienert[hidden]
charm137[3 comments hidden]
So they clearly have a product, a strategy around it and perhaps the compliance scaffolding (SOC2 Type II etc) that may be needed before actually being able to charge money for it. They also have the right brand names associated with the founding team. Execution, so far, seems good enough to create a splash, at least.
As an investor, the question(s) to ask is (in my view): "How do they make money? Will that way to make money survive?". The answer to the first: selling input tokens and perhaps subscriptions/credits eventually.The answer to the second: "Yes, but with the risk of unit revenues declining faster than their unit costs". How can they mitigate this problem: by being big (scale / mindshare etc) so that their unit costs (including for customer acq) fall faster than their unit revenues will - I believe that is the question most AI companies are trying to tackle these days. Any new competitor will have to tackle basic fixed costs (of getting started) first before even getting to the stage of having the luxury of worrying about unit-economics.
So yes, they might eventually be competed away but whoever is in their team is trying hard to make a useful product/ecosystem and that should be applauded, not ridiculed with "it's all marketing". This is way more than a simple github/huggingface-based open-source replica solution can hope to achieve without institutional backing (either big-tech or system-integrators).
What should rightly be questioned, of course, are the valuations the VCs are providing to them in hopes of passing this hot potato to a willing buyer (say a hardware maker like NVidia) - the incentives there are very well defined and depend very much on perceived TAM (which lately is on very shaky ground given how far token pricing has fallen causing, among other things, OpenAI to "miss" on the market's expectations for annualized revenues, even before they're listed!) [1]
[1]: https://www.ft.com/content/b66a9858-f8fb-46cb-b506-44bfe26fc...
notfromhere[hidden]
9dev[hidden]
So yeah, not too much enterprise sales experience there for sure.
neilellis[2 comments hidden]
phoghed[hidden]
Then everyone breathlessly repeats this story about how Jev is useless because open source models they never tried claim to do the same thing and better.
fastball[hidden]
jofzar[hidden]
vrighter[hidden]
Ozzie_osman[2 comments hidden]
c7b[hidden]
fennecbutt[2 comments hidden]
vrganj[hidden]
> Fictitious capital could be defined as a capitalisation on property ownership. Such ownership is real and legally enforced, as are the profits made from it, but the capital involved is fictitious; it is "money that is thrown into circulation as capital without any material basis in commodities or productive activivity".
https://en.wikipedia.org/wiki/Fictitious_capital