Alignment Forum
Alignment Forum publishes articles covering AI, LLM. A trusted source for AI and technology insights.
Profile generated by AI for Anything

The Long (Self-)Correction
3h

Challenge: Hand coding weights for efficient sequence memorisation
1d

Analysis: Score-seeking misalignment in the OpenAI–Hugging Face incident and its existential risk implications
1d
![[Paper] Stringological sequence prediction II](https://res.cloudinary.com/lesswrong-2-0/image/upload/v1654295382/new_mississippi_river_fjdmww.jpg)
[Paper] Stringological sequence prediction II
2d

Towards surfacing model algorithms with meta-tokens in the J-Space
4d
A Red Line and Oversight Framework for Government AI Contracts
6d

Endogenous Alignment
6d

Should we benchmark conceptual capabilities using judgment prediction tasks?
7d

Announcing the Corrigibility Research Fund
7d

Why I Left Google DeepMind
9d
Open Distillation of Hereditary Traits
10d

The original title is "Prism: Automating Science-of-Evals Research"
11d

Independent alignment of language models
12d
From wantons to moral agents
12d
The current bottleneck is political will, not research
13d

Value generalisation: value correction
14d
The original title is a question: "How robust are natural language autoencoders to initialization?"
14d
AI 2040: Plan A
15d

Announcing our $160M grant from Coefficient Giving
15d
Modular Pretraining Enables Access Control
16d

Notes on technical alignment via human-like social drives
16d

Data filtering works a lot worse than you would expect
17d

Pragmatic FDT, and predictors as game theory
21d
What Capable Agents Must Know: Why AI Consciousness May Be an Inevitable Byproduct of Capability
24d