Ai Alignment
78%14 posts analyzed over the last 12 weeks
Over the last 12 weeks
Average across all posts
This month vs. previous
14 posts analyzed over the last 12 weeks
Over the last 12 weeks
Average across all posts
This month vs. previous
Growth slope (6 weeks)
Best week: 27 avr. (28 avg. likes)
Sairam Sundaresan
AI Engineering Leader | Author of AI for the Rest of Us | I help engineers land AI roles and companies build valuable products
Most AI engineers fine-tune models The top 1% fine-tune themselves with these 37. 🔸 2017 - 2019: The Foundations (The Transformer Revolution Begins) 1. Attention Is All You Need: https://lnkd.in/gHY8rjdw 2. BERT: htt…
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
Everyone is arguing about which path to superintelligence is fastest. That is the wrong argument. DeepMind just published a careful 57 page map of how we get from AGI to ASI. It lays out four routes: 1. Scaling: keep …
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
To predict a process well, a model has to build an internal model of whatever generates it. That sounds like hand-waving until you can point at where it lives. 'Transformers represent belief state geometry in their res…
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
We are about to hand AI agents real decisions, and most of them cannot prove who they are. I have spent the past year building identity and audit infrastructure for AI agents, and this is the gap I keep coming back to. …
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
Most worries about using AI to automate alignment research focus on a scheming agent that sabotages the work on purpose. A new paper from Bowkis, Buhl, Pfau, and Irving (arXiv 2605.06390) argues the more likely failure …
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
Forty researchers from OpenAI, Google DeepMind, Anthropic and Meta agreed on one thing. We might be about to lose the ability to read what our models are thinking. That is the claim in "Chain of Thought Monitorability"…
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
Honesty in a model is not a vibe you sense in the tone. It can be defined. And once defined, measured. A lie is asserting what you believe to be false. So deception becomes a quantity: the gap between what a model sta…
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
Loss going down tells you almost nothing about what a model learned. The number falls smoothly. The thing underneath it does not. Singular learning theory treats a network as a singular model, and the loss landscape st…
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
Most AI oversight stops at the software layer. That is the quiet weakness almost nobody states out loud. If the logs can be edited by whoever controls the host, oversight is a story told after the fact by the party you…
Pavan Kumar Dubasi
AI Safety & Alignment Engineer | LLM Fine-tuning, RLHF, Red-Teaming & AI Governance | EU AI Act Compliance | Multi-Agent Systems | W3C Verifiable Credentials | IEEE Published Researcher | Founder, VibeTensor
A year ago, almost every safety paper I read was about the weights. Now the centre of gravity has moved to the trajectory: not what the model is, but what the agent actually did, step by step. I track this landscape cl…