Hacker Newsnew | past | comments | ask | show | jobs | submit | fabmilo's commentslogin

Location: San Francisco, CA, USA Remote: Yes Willing to relocate: No

Technologies: Python, Go, TypeScript, PyTorch

Linkedin: http://www.linkedin.com/in/fabmilo

Hi, I’m Fabrizio Milo, a senior AI/ML engineer, large-scale systems architect, and former technical co-founder. I’ve spent my career building production AI, ML infrastructure, and high-scale backend systems across startups and growth-stage companies.

Most recently, I’ve been building AI platforms for LLM training and fine-tuning, RAG databases, local and remote LLM inference, semantic retrieval, and agentic orchestration for business intelligence and code generation. I’ve also contributed to open source ML projects including GPT-Neo, TensorFlow and published research on synthetic data from LLMs.

Previously, I was VP of Technology at ZELIG, where I led the virtual try-on AI research roadmap and managed a cross-functional team building ML/3D systems for fashion retail. Before that, I was Head of Machine Learning Engineering at Recurrency, where I hired and led a 6-person ML/platform team and shipped demand forecasting, dynamic pricing, and recommendation systems on AWS/Snowflake/SageMaker. I also co-founded Passio, where I built the technical foundation for an on-device Nutrition-AI SDK with real-time computer vision inference.

Earlier in my career, I built scalable systems at TheRealReal and Scopely, optimized CUDA kernels at NVIDIA, and worked on real-time market-data and high-performance systems. I’m strongest where AI research, production engineering, and startup execution meet: taking ambiguous technical/product goals and turning them into shipped systems, teams, and infrastructure.

I can architect and build anything you need given enough compute and time.


Writing it thinking. We developed our brain together with our hands. It feels slow but is actually faster for the end goal.


I am fascinated by this example of using AI to improve AI. I won a small prize using this technique on helion kernels at a pytorch hackathon in SF.

The next step are: - give the agent the whole deep learning literature research and do tree search over the various ideas that have been proposed in the past. - have some distributed notepad that any of these agents can read and improve upon.


Was thinking the same thing. probably once a day would be more than enough. if you really want a minute by minute probably a delta file from the previous day should be more than enough.


indeed. make a loom showing us why is better.


There is tons of good advice. This blog post can be easily turned into a skill for agents.


This is surprising to me. The advice about what team members should be able to is the stuff I find agents least capable of doing, e.g. autonomously identifying the most important work and knowing when something is done.


New generative modeling using a single inference step


Very impressive work from Waymo. The driving with a tornado in the horizon example kind of struck my imagination, many people actually panic in such scenarios. I wonder though the compute requirements to run these simulations and producing so many data points.


because of the principle: you only understand what you can create. You think you know something until you have to re-create it from scratch.


VAE for real time video generation, WAN 2.1 / Matrix Game 2.0


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: