Hacker Newsnew | past | comments | ask | show | jobs | submit | hleszek's commentslogin

That benchmark could really be a good AGI test. Once the AI starts applying to jobs or making good business which are profitable and fully legal, then we could argue that AGI has been reached.

If you could construct a sandbox to test this, where it doesn’t touch the real economy, then yeah it’s a great benchmark.

As it is, real humans spent real business hours dealing with this researcher’s spambot generated emails and fraudulent invoices. Individual recipients reported feeling harassed.

This isn’t a good benchmark. It’s a series of socially destructive crimes committed by the researchers and then documented and published on the internet.


> Once the AI starts applying to jobs

I guess that's already a reality? [1-2]

[1] https://github.com/jaimaann/LangHire

[2] https://github.com/adrianhajdin/job_pilot

(among many other similar projects)


Seems beyond AGI at that point? Most humans wouldn't be able to make a good business that is profitable.

Sure, but you can do that at home too.

Earth is home.

Ask the AI to create a detailed spec according to a few simple requirements. Review the spec yourself and correct what you want changed. Then ask the AI to implement the spec. Each time you request something new, ask the AI to update the spec as well.


This so much more annoying and circuitous than writing code.


Then just write the code, my man! This guy in this post is CHOOSING to use the LLM.

Choose what makes you happy!


It isn't because I'm doing something else while it churns away.


I just completed the course and a big part of it was in fact to warn not to use it to send private data.


No, it works using Malta eID identification system which needs you to be a citizen or a resident.


Why are tiles small BTW? Could we use tiles as big as normal solar panels?


My guess would be they are the same size as shingle strips, to make it easier to work with for regular installers rather than specialists.

These things carry a lot of current though, so I would certainly not trust anyone without proper tools and training to put them on a roof.


It is not a meme, it's an xkcd: https://xkcd.com/810/


Like the Delamain AI in Cyberpunk. You would need to allow anonymous payments with cryptocurrencies for that, but it's coming for sure.


For open-weights models, censorship removal is now a "solved" problem. If you wait a few days after a new model release, someone will have made a heretic ( https://github.com/p-e-w/heretic ) version with the censorship removed, so in a way the only use for censorship now is to avoid lawsuits, not reduce improper usage.


Any time I've tried an "abliterated" model, heretic or other, it has always damaged the capabilities of the original model and will still often refuse or produce garbage at a lot of "unsafe" requests.


Abliteration can't teach the model something that wasn't in pre-training, it's just fixing refusals from post-training. I don't find the delta to be that big in practice and it really depends on what you're doing with the models anyway. If your primary usecase is sexy roleplay I think the loss of absolute capability is probably worth the abliteration, for malware research it's probably better to just jailbreak.

I've mostly found that finetunes and abliterations are of limited use but that's recently changed for me. My default model for the past week or so has been a Qwen 3.6 tuned on Opus 4.7, it's definitely a bit worse than the base Qwen in terms of precision and "intelligence", but it MORE than makes up for it in response style. Way easier to get it to write things that I want to read, it's way more terse, way fewer emoji. Best local rubber duck by far.


Do you have a hugging face link?



There are many abliterations which work quite well. Older techniques do suffer from quality issues, but more recent ones do a much better job. In particular, the older approaches did poorly on MoE models.

Another likely problem you're running into: the problems with older techniques compound with quantization. Anything less than 5-bit quant is going to give you some pretty sketchy outputs, in my experience.


The problem is the heretic and abliteration versions are dog shit quality compared to the non-edited versions and much more likely to hallucinate.

AFAIK abliteration without quality reduction isn’t even possible without some quality reduction, even if it’s marginal. All the benchmarks reflect this.


Please mention and support llama.cpp directly instead of ollama.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: