r/MachineLearning
r/MachineLearning is a Reddit community for machine learning researchers and enthusiasts. It features discussions on topics like networking at conferences and the long-term value of AI research.
Profile generated by AI for Anything
![Loss functions in Instance Representation Learning [R]](https://preview.redd.it/3l7mtxoc3bah1.png?width=140&height=27&auto=webp&s=8426b12f6ec1f44b193529124dee890e0642ad25)
Loss functions in Instance Representation Learning [R]
13d
![I'm trying to implement CALM paper, and I have some questions. [P]](https://preview.redd.it/kr4u22yfx8ah1.png?width=140&height=83&auto=webp&s=784c46c82400e669571b4d8a7dcdc997ad0fba57)
I'm trying to implement CALM paper, and I have some questions. [P]
13d

The original title is about a quiz that tells you which LLM you align with most. Let me rewrite this for a mobile feed.
14d
![NagaTranslate: Building a translation and voice pipeline for low-resource Nagaland creoles (Whisper, VITS, LLMs) [P]](https://preview.redd.it/bu6xsk4hvx9h1.jpg?width=140&height=63&auto=webp&s=0a4c589616d351d10d735940874706494e48d408)
NagaTranslate: Building a translation and voice pipeline for low-resource Nagaland creoles (Whisper, VITS, LLMs) [P]
15d

Attention pathologies stem from norm
17d
![[R] Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost](https://external-preview.redd.it/q3evP6JeDpAC2MdSQHWYxnCYTqbJkElIQsLFqVSdkss.png?width=640&crop=smart&auto=webp&s=de730fbf7ecace6df0036b21470c16a2d4feacfb)
[R] Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost
17d
![[R] All Routes Lead to Collapse: attention sinks, representation collapse, and norm stratification are what content-based routing does under a norm-blind metric](https://external-preview.redd.it/q3evP6JeDpAC2MdSQHWYxnCYTqbJkElIQsLFqVSdkss.png?width=640&crop=smart&auto=webp&s=de730fbf7ecace6df0036b21470c16a2d4feacfb)
[R] All Routes Lead to Collapse: attention sinks, representation collapse, and norm stratification are what content-based routing does under a norm-blind metric
17d
![CALHippo - Mapping neurons and glial cells in the human brain hippocampus in 3D using SOTA segmentation and density estimation models [R]](https://preview.redd.it/m8eyacfmbf9h1.gif?width=640&crop=smart&s=1a9d654de34977e02d4c3b3a30f0f9e2d36a5c35)
CALHippo - Mapping neurons and glial cells in the human brain hippocampus in 3D using SOTA segmentation and density estimation models [R]
18d
![High Dimensional, Dynamic Rotary Positional Embedding [P]](https://external-preview.redd.it/Go7zlxhewkLxNN5-ZvZe623w5Zrdi3SXYEIr0JeEGQk.png?width=140&height=75&auto=webp&s=2d3a7ad647024e077a4b7f7b5746c806eba71b8a)
High Dimensional, Dynamic Rotary Positional Embedding [P]
18d
![I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]](https://preview.redd.it/4vj50mvhu79h1.png?width=140&height=63&auto=webp&s=f53b566e7aa9a25215aa77fcf3ed0b16e426e2a1)
I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]
19d
![DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]](https://preview.redd.it/lacvagyr159h1.png?width=140&height=89&auto=webp&s=14f97a97511fbfe2fd767e4dc986ce0b4da5c73e)
DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]
19d

Free lecture covers multivariate probability
25d
![Next-Latent Prediction Transformers [R]](https://preview.redd.it/efm7zazr2t7h1.png?width=140&height=90&auto=webp&s=c1b7070ca3de62bdc276d7a185c72f6737e6f92e)
Next-Latent Prediction Transformers [R]
26d
![What is Speculative Decoding? (trending on paperswithco.de) [R]](https://preview.redd.it/dm4nh4t71o7h1.png?width=140&height=90&auto=webp&s=4bde95d9237d3d2f4f1139976ad15967ef1f3f5c)
What is Speculative Decoding? (trending on paperswithco.de) [R]
26d
![AI language models have favorite names, and we mapped them [R]](https://external-preview.redd.it/q3evP6JeDpAC2MdSQHWYxnCYTqbJkElIQsLFqVSdkss.png?width=640&crop=smart&auto=webp&s=de730fbf7ecace6df0036b21470c16a2d4feacfb)
AI language models have favorite names, and we mapped them [R]
27d
![go from zero to claude code pro in one day — hands on bootcamp may 30 [D]](https://preview.redd.it/3cwac3a4ig0h1.png?width=140&height=73&auto=webp&s=926b8295bde8ff17f14d9d7be03e9b9a048a6a12)
go from zero to claude code pro in one day — hands on bootcamp may 30 [D]
63d
![Signals: finding the most informative agent traces without LLM judges [R]](https://preview.redd.it/nauai52sgc0h1.png?width=640&crop=smart&auto=webp&s=b9a1d06b2ba6b8e05d6ef0c125f39510f7e0806b)
Signals: finding the most informative agent traces without LLM judges [R]
63d
![Steam Simularity Reccomender Student Project [p]](https://preview.redd.it/kyevk5w4h70h1.png?width=140&height=138&auto=webp&s=d8a68dd3d8ca8f7fc2df6ede9b0869b4042b861f)
Steam Simularity Reccomender Student Project [p]
64d
![Steam Similarity Recommender Find your next favorite game and learn WHY (student project)[P]](https://preview.redd.it/st5zcvr0lzzg1.png?width=140&height=128&auto=webp&s=2475c53ab4705feda695bc748d39aa668bb81afe)
Steam Similarity Recommender Find your next favorite game and learn WHY (student project)[P]
65d
![Formalizing statistical learning theory in Lean 4 [R]](https://external-preview.redd.it/vcCPMFQJ_Ba1RGnmGbl8xPO9fb1vZ01Ht3Hg4xF7i0Q.png?width=640&crop=smart&auto=webp&s=66c23bc95d8b7a4cf205d906d8b6fed9fda955ab)
Formalizing statistical learning theory in Lean 4 [R]
65d
![What should a PyTorch training end-of-run performance summary show? [D]](https://preview.redd.it/2q71s9ltkvzg1.png?width=140&height=123&auto=webp&s=77945925d0fb80c62d90ad917fad64c3af772f43)
What should a PyTorch training end-of-run performance summary show? [D]
66d
![Steam Similarity Recommender [P]](https://preview.redd.it/njoqyt939uzg1.png?width=140&height=128&auto=webp&s=1511dec07fceeeb727b5c1219bbc1e10017a1918)
Steam Similarity Recommender [P]
66d
![Transformer Math Explorer [P]](https://external-preview.redd.it/BFIg4KJ_Lb1cT_fEp_aR5T8A2tbOhZKO3DwQFn9Xxf8.png?width=640&crop=smart&auto=webp&s=cae1df99908e4bdf90f43d52052ca6c3ff4938ea)
Transformer Math Explorer [P]
66d
![Question about PLS-DA hyperparameter tuning [R]](https://preview.redd.it/701knkltbdzg1.png?width=140&height=86&auto=webp&s=be7986f3d4908ceaf4f9969b26a3eafd1012a3e5)
Question about PLS-DA hyperparameter tuning [R]
68d