Just wanted to share my review and experiences of using Lumo 2.0, which is Proton's new LLM.
Given that many of us use Proton services and are privacy conscious, I thought this might be of interest.
The below is a lightly edited from another post.
It broadly speaks to the integration (or lack thereof) of Lumo with other Proton services, rather than the privacy angle (which I am sure you're aware of), though I can speak to that too. EDIT: see my below follow up.
I see no reason at this point to continue with my subscription, and if anything, I'm thankful for the shot in the arm to improve my self hosted stack / create an equivalent or better service for myself.
TL;DR: With some elbow grease and know how you can probably create something equivalent or better home.
I keep seeing suggestions here that Proton should adopt Kimi K3, or otherwise move to a much larger and more fashionable model.
I am increasingly unconvinced that model capability is Lumo’s primary problem.
Lumo Lite appears very likely to be based on a Qwen 27B-class model, judging from the behaviour and the Artificial Analysis results Proton have stated.
(That identification is not proven, obviously, but it is plausible enough for my broader point).
A model in that class should already be an excellent all-round foundation for Lumo.
My humble suggestion:
Serve it at FP8 or comparable precision. Optimise it properly. Keep latency reasonable. Increase the context length.
Then spend the engineering effort on the things that would actually distinguish Lumo from every other chatbot:
-
Reliable Proton Mail, Drive, Calendar and Contacts integration
-
Authoritative and current Proton product information
-
Deterministic tools for calculations, dates, prices, plan limits and account features
-
Exact-span retrieval rather than loose paraphrasing for proton_info
-
Visible source dates and document versions
-
A clear distinction between retrieved evidence and the model’s interpretation.
At present, Lumo can (apparently) retrieve information about Proton’s own products and then produce an answer that reverses or mangles the source.
Something like “it is not $12.99” becoming “$12.99” is not a problem that Kimi K3 or GLM inherently solves.
It is a grounding and validation problem.
A larger model may paper over some failures. It may phrase uncertainty more elegantly. It may recover from poor retrieval more often.
But it does not make stale information current, and it does not make an unreliable tool pipeline deterministic.
IMESHO, Proton’s strategic advantage is not that it can host the largest open model.
Other companies will always have larger models, more compute and faster release cycles.
Its advantage is that it owns a private productivity ecosystem, with only one real mind share competitor.
A smaller model that can reliably search my mail, locate a Drive document, understand my calendar, retrieve the correct Proton support information, chat, OCR (native ability with latter Qwen models iirc), image manipulation and perform bounded actions would be considerably more useful than a stronger general-purpose chatbot with shallow integration.
I would much rather Proton solidify Lumo Lite into a dependable, deeply integrated assistant than keep changing the model underneath it in pursuit of benchmark gains.
The model is probably already good enough.
The product around it is not.
ICBW and YMMV.
Proton can't read your mail (big "*"); it's decrypted client-side. Feeding it to their LLM would violate that. I'm unsure about their other services.