Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

LLaMA2 is English-centric fwiw, censored for safety, and black-box on its contents. Happy to be proved wrong


Think of the censoring as comedy.

It doesn’t like Kanye music and will tell you to listen to the song ‘Happy’ instead.

That’s pretty hilariously dystopian and therefore quite funny.


You're right, none of that is disputed though?


On the second point, there's an uncensored version available and assorted trained-for-purpose derivatives of it.


Are these uncensored or decensored? If rlhf removes intelligence at any rate, I wouldn’t expect that intelligence to come back with a tune that’s let’s it say curse words and talk about religions


The foundation model is not “censored”, it’s the RLHF’d “chat” versions that are. Meta released both.


They are very much uncensored, give it a try yourself: `ollama run llama2-uncensored` [0]

It will be happy to curse, talk about religions, help you cook illicit substances or do all sorts of other stuff.

[0] https://ollama.ai/


links please!



These are not the uncensored models. It's a fine tune of the censored models [1].

> Filter refusals and bias from the dataset -> finetune the model -> release.

The alignment tax should still exist, maybe doubly so.

[1] https://erichartford.com/uncensored-models#heading-lets-get-...


The base model is uncensored. From a glance at your link, that is about uncensoring the conversational training set.


I don't think so: https://erichartford.com/uncensored-models#heading-whats-an-...

> Most of these models (for example, Alpaca, Vicuna, WizardLM, MPT-7B-Chat, Wizard-Vicuna, GPT4-X-Vicuna) have some sort of embedded alignment

> The reason these models are aligned is that they are trained with data that was generated by ChatGPT, which itself is aligned by an alignment team at OpenAI.


Most of those are fine-tunes of the base model. The fine-tuning data is 'aligned'. The uncensored fine-tune training data is edited to remove the "I can't help you with that" responses.


Yes, as stated in my earlier comment: there's an alignment tax, and then almost certainly an un-alignment tax, on top of that, compared to the raw, unaligned/uncensored, models.


that is interesting, but the "un-training" shown by EricH. is simply re-running some fine-tuning on the same public base model, regarding "refusals".. and it is expen$ive to do that, too.


Llama2 based models are actually quite proficient in at least French and German, if not more western european languages.


well, your username suggests you'd enjoy Mistral 7b more /jk




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: