Rendered at 16:38:55 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Whitespace 1 days ago [-]
I have a weightlifting spreadsheet with weight on the vertical axis and reps on the horizontal axis. The value of each cell is the estimated 1 rep max if I accomplish that lift. In theory if my e1RM is 100kg then I can lift any permutation of (weight,reps) that have the same e1RM. This is akin to knowing Pareto Frontier of my current strength.
I use conditional formatting to color cells according to the probability that I can lift them—if I lifted 50kg for 10 reps then I can definitely do 50kg for 9 reps, so that cell is green. But if e1RM(50,10) > e1RM(40,15) then I can probably do that too so it's light green. The visualization naturally becomes Pareto-like.
If I'm feeling strong I can aim for higher weight, lower reps. Or if I'm feeling weak I can close out a (weight, reps) that's below my current e1RM but I haven't accomplished yet. The end result is that I'm always "accomplishing" some sort of PR no matter how I feel.
I call this e1RM Bingo.
jerkstate 1 days ago [-]
I wrote this app as a SPA! It uses a curve formulation similar to Brzycki, except I added a “shape” parameter (an exponent gamma between 0 and 1) that slopes the 1rm downwards at the right side.
My main finding for “pick whatever weight you want today” was that picking a lot of different weights made the curve less identifiable, so my latest iteration encourages you to pick a ladder for a few sentinel exercises per mesocycle in order to improve the statistical power. In addition, strength improves more quickly at >80% of 1RM, and hypertrophy depends on proximity to failure, so if you pick a lower weight, you really need to go to failure, which burns you out for the rest of your session, where leaving 1-2 reps in reserve is probably sufficient for hypertrophy and leaves a lot more gas in the tank for the rest of the session. Definitely open to suggestion/discussion here.
https://curvefit.app (it runs on Cloudflare free tier, so I won’t have to start running ads or charging until I hit a couple thousand users)
fudged71 19 hours ago [-]
This is phenomenal, I'm definitely going to try this. Any chance this is OSS or plans to publish in the future?
jerkstate 18 hours ago [-]
There’s no particular reason it’s not OSS, but my main interest is collecting a lot of data on different athletes and publishing original research. Most weightlifting studies are small n and over a short amount of time. My particular interest is how volume, load, and fatigue are related to strength, endurance, and compliance over time. My intention is to run it for a while, look at the data to generate some hypotheses, pre-register them, then run some experiments (and by that I mean just keep collecting data). If someone else was particularly interested in this goal, I would definitely invite them to the project. That’s why it was important for me to design it to be hosted for just the cost of the domain name, because I don’t really intend to make money from it, I’m just interested in the data.
747-8I 21 hours ago [-]
Great - commenting to refer to this
cman1444 1 days ago [-]
Could you please share this spreadsheet? I would really love to have my own version of this.
kachnuv_ocasek 1 days ago [-]
Just copy-paste that description to Claude and have it create the spreadsheet.
Whitespace 18 hours ago [-]
The last time I tried this was back with Opus 4.6, and it was ok. I tried it with Fable 5 High just now and I was very impressed with the output. It took 7 minutes and one turn.
I'm not one to believe in all the one-shot hype, but this was pretty good.
joncrane 1 days ago [-]
This is a cool way to gamify weightlifting. Cheers!
godwinson__4-8 1 days ago [-]
Indeed, GP should take a spin at turning into an app. Could be worthwhile to have Claude take a first stab at a MVP.
If pursued, good luck!
jerkstate 1 days ago [-]
If this is something you are interested in, I did make a mobile friendly SPA similar to this: https://curvefit.app
xnx 20 hours ago [-]
Clever name
wollowollo 1 days ago [-]
Respectfully, that's a cool illustration of the idea of xRMs etc but is missing the whole point of programming for higher or lower reps. E.g. lower reps are more stressful / higher cost of recovery but more strength-specific; high reps are better for hypertrophy work. But then, any well designed program will have you working across a range of rep ranges and so on.
Please don't make an app based on this.
jerkstate 1 days ago [-]
> high reps are better for hypertrophy work
Some nuance here: the latest research shows that proximity to failure is the main hypertrophy driver regardless of load and rep count; high rep count makes proximity to failure harder to gauge; so high load/low reps close to failure is probably better for hypertrophy (there are other good reasons to do higher reps/lower load work though)
bob1029 1 days ago [-]
The most effective (difficult) training regimens usually avoid the middle of the distribution. You generally want to be operating at the extremes with some rotation schedule or duty cycle. High intensity interval training is an example of this philosophy that occurs within a single workout session.
If you want the most 'optimal' form of this (aka, hell on earth), you should purchase a rowing machine. Being able to engage with very aggressive, full-body exercise every single day without exceptions is almost like cheating biology. You can maintain a 2-3x VO2 max premium over your peers with very little risk of injury.
21 hours ago [-]
MSKJ 23 hours ago [-]
Respectfully, that's missing the point of the comment. It's a fun thing to hit PRs, not everything needs a 'well actually'
deadbabe 24 hours ago [-]
Respectfully, it’s nothing new. Weightlifting industry has known this concept forever, it’s often just expressed as charts rather than graphs, as it is easier to interpret.
But they go even a step further, they extend into 3 dimensions to also add body weight as a variable. So your graph would really have to be a 3D volume. Because different levels of body weight have different capabilities.
rafabulsing 17 hours ago [-]
Respectfully, his graph does not need 3 dimensions because it's a personal spreadsheet he uses just for his own training, so he can just display the data for his exact body weight.
Aachen 1 days ago [-]
Misread the title and got excited about a Pareto font, that is, the best possible font (presumably: distinct l/I, O/0, scores within error margins of the top readability and reading speed scores, widely available, etc.)
Maybe in vein but did anyone already figure this one out? The closest I got was PT sans, open-licensed commissioned by the Russian ministry for communication (I found it surprising that a country that doesn't use Latin script made the best font!), but it's not widely shipped so you need to figure out how to include font files whenever you want to use it
gadders 1 days ago [-]
I thought it was some sort of Italian activism group.
"The Pareto Front today claimed responsiblity for...."
lucaslazarus 23 hours ago [-]
You mean the People’s front of Pareto!
croisillon 21 hours ago [-]
splitters!
orthoxerox 1 days ago [-]
...20% of the attacks causing 80% of the casualties?
mdnahas 1 days ago [-]
As someone with an Italian grandmother and both a CS and Econ degree, I got a great laugh out of this joke! Bravo!
pphysch 1 days ago [-]
No!! That is the Front of Pareto[1], a totally different group. Our Pareto Front only causes 20% of the casualties.
I didn't know fonts can have options. Another learning curve on how to enable that in Latex/Html/Libreoffice/anywhere else I use fonts ^^'. But still helpful to know about!
ss02 disambiguation seems to be the one I'd be wanting to turn on, with tnum for monospace numbers being a good option as well that I hadn't even realised I wanted from a font!
airstrike 24 hours ago [-]
Yeah, those features are awesome! Sadly support for them is quite lacking in applications...
Tabular numbers are awesome!
notpushkin 22 hours ago [-]
I feel like there should be a GUI tool to bake in any alternates into the font itself.
look at Atkinson Hyperlegible? commissioned by the Braille foundation for low-vision readers which means it's very readable
Aachen 22 hours ago [-]
Huh, I'm positive this was one of the first I considered (or perhaps it was a different dyslexia font) but I can't find the reason why I would have preferred PT Sans over this. Each has some pros and cons: PT Sans has same-width digits; Hyperlegible has a nicer-looking Q (imo); PT Sans has a normal-looking zero (not struck); Hyperlegible has a normal-looking 'fi' combination (no ligature that merges the two characters into a new one); Hyperlegible has a 6 distinct from an inverted 9; PT Sans needs a bit less horizontal space. The font specifically designed for accessibility should be the easy winner if all else is equal anyway
They apparently released Hyperlegible Next in 2025 which, flipping between tabs on Google Fonts (since the original website doesn't show the fonts), is nearly identical but has five new weight settings (nobody should imo ever use thin fonts though, it noticeably harms readability for me and my sight is only the tiniest bit below normal vision, but ok it's an option) and improved kerning (the original font had extremely little space between 'll', for example)
The 2025 version sadly doesn't ship with my version of TexLive, but the original (from 2020) already does so that makes it easy to use as well! Cool stuff, thanks for the tip :)
adornKey 1 days ago [-]
I'd also be interested in the Pareto Front of Fonts. That would be the final font collection - to rule them all.
arduanika 1 days ago [-]
No, it's far from the best possible font, but to its credit, it gets most aspects of typography right by just focusing on the ~1/5 of the requirements that actually really count.
Aachen 23 hours ago [-]
What are the 4/5ths that you don't like, and which one(s) would you recommend instead? (I'm mostly interested/focussed on the sans variant but also happy to hear about serif and mono)
arduanika 22 hours ago [-]
It's not that I don't like them. It's more just that some of these design considerations are not worth the effort to fret over, and if you're smart about your time, doing 1/5 of the work will get you something like ~4/5 of the results.
notpushkin 22 hours ago [-]
...so it is the Pareto font indeed!
joshka 1 days ago [-]
lol same :D
CodeIsTheEnd 1 days ago [-]
I am training for a marathon, and, as I increase both by distance and pace, I am always excited when I have a "Pareto run": a run along the Pareto frontier of me trying to maximize distance and speed.
When explaining it to some coworkers, I stumbled on a fairly intuitive explanation: "I've run farther before, and I've run faster before, but I've never run _this_ far, _this fast."
There was some pushback about why not just call it a PR (personal record), but I would only use that term for fixed distances (1mi, 5k, 10k, etc.) or a consistent route that I've run many times before. Nobody would say "I set my 7.40 mile PR today." More importantly, it misses the comparison to all farther (and faster) runs—it's not exciting to set a 5k PR just because you've barely run that distance before, and the pace is actually slower that a 10k you've done.
(Had a Pareto run of 7.40 miles @ 6:28/mi last week!)
froxtrot 1 days ago [-]
Not relevant to pareto, but that's a really fun way to look at running. Not quite as fast as your pareto run shows, but I'll definitely keep that metric back of mind to keep the psyche high for running.
reedf1 1 days ago [-]
The cycling equivalent is your power curve, i.e. the longest you've held a power for a certain time interval.
voidhorse 15 hours ago [-]
So, while it's true that your runs with high speed and distance when both are considered are Pareto points, your max speed run and max distance run overall are also Pareto points. So calling these high distance+speed runs "Pareto" doesn't actually distinguish them completely from other runs.
A point is Pareto so long as it is non-dominated—that is, you're not looking for dominating points, you're looking for points that are "no worse" than all others, in all criteria, when you consider that point as a reference.
So your Pareto Runs are indeed Pareto points. However, your run with your fastest possible speed, even if your distance was really bad, is also still a Pareto efficient point.
(Taking >= as more efficient here) By definition, the point A is Pareto if there is no point B such that in all criteria, B >= A, and for at least one criteria B > A. Take the run with the best speed. It is Pareto because we cannot find a single point B that satisfies both of these conditions. Your "Pareto Run" doesn't satisfy this set of conditions because it is worse in terms of speed, even if it has better distance than the max speed point.
The only way your Pareto runs would be the only Pareto points in your record is if they simultaneously hit maxima for distance and speed when compared to all other points. So, for them to be the sole Pareto point, the clause ""I've run farther before, and I've run faster before..." would have to be false! The point would have to break both your all time records to be the solitary Pareto point. With running, because of how speed and distance are related this will basically never happen.
The definition of Pareto efficiency is essentially negative in nature--it's not about finding specific dominating points, it's about finding points that are not dominated by any others on any criterion, period. All criteria are weighted equally in the search for Pareto points. It doesn't build in any weighting like considering maximum across criteria as "better" than points that only maximize one criteria. For a "biobjective" problem like your runs, the Pareto set will always contain the points (MAX, -) and (-, MAX)--they may not be unique over the criteria but there will always be at least one representative for each, I believe.
speedstyle 8 hours ago [-]
They never said it was the only Pareto run? Just a new one, improving the overall Pareto front
voidhorse 3 hours ago [-]
Ugh, you're right! I read too quickly. Please disregard me OP. I'll leave the comment here just in case it clarifies things for poor readers such as myself lol.
ChatGPT 5.6 Luna on the right (cheaper) cover most of the frontier, with a point for Deepseek flash, and higher performance overlapping heavily between 5.6 Sol and Fable.
That DeepSeek point will probably move back towards Luna as deepseek announced a "significant" price increase coming to their API [1], which kind of demonstrates that beating the Pareto frontier is where the difficulty actually is).
I've been wondering if OpenAI make Luna artificially cheap to get people into their eco system.
I think it's great and hope the price can stay the same.
bob1029 1 days ago [-]
I think Luna might be just small enough to provide some kind of stepwise improvement in how it is hosted.
Going from 81GB of weights to 79GB of weights can mean a 50% reduction in GPU capacity required.
If you can fit a model in just one GPU (or rack) as opposed to across an entire datacenter, the latency gains can be substantial too. If you can reduce token latency by half, that would double the amount of customers you could support.
kazinator 2 hours ago [-]
A kind of Pareto domination criterion is used in C++ for determining overload resolution: which function overload gets the call.
The objectives are matching arguments to parameters.
A set of functions is identified among the candidates: those that are possible for the call at all, like having a compatible number of parameters.
Essentially, the overload rule says that the Pareto front set of candidates must contain one member, otherwise the call is considered ambiguous, and diagnosable rule violation.
The objectives being optimized are individual parameter positions, each in the dimension of suitability: being a better match.
One candidate is better than another if it is no worse a type match in every parameter, and strictly better in at least one parameter.
bob1029 1 days ago [-]
Pareto front sounds like an interesting way to optimize, but it suffers from the curse of dimensionality just like anything else.
As the number of objectives (dimensions) increases, the number of samples you need to cover the frontier increases exponentially. You will very rarely find solutions that actually dominate other solutions in many practical optimization scenarios. With 2 dimensions you have a 25% chance of domination. With 10 dimensions it's a .098% chance.
The most useful cases I've seen tend to occur where we just optimize for two things at once. The chances of domination are high, it's easy to visualize and very efficient to implement. As we get into higher dimensional spaces, things get weird really fast.
peri-cl 1 days ago [-]
> "As we get into higher dimensional spaces, things get weird really fast."
The geometric problem of computing a d-dimensional Pareto set of cardinality n
has a truly weird property not covered by the computational complexity discussion on that page. It says there's an algorithm achieving O(n log(n)^(d-3) log log n), which is true and also a lie. The algorithm that achieves that asymptotic form is a galactic algorithm; and not an ordinary one in the sense of "has a large constant multiplicative factor", but one with this property (I've never found any other algorithm which exhibits it):
The runtime is within a bounded constant factor of n^2, for all n up to some critical N whose size is exponential in d (I think it was exactly 2^d or something).
I.e. the runtime has "two shapes": it's purely quadratic up to a galactically-large constant, and thereafter has a transition into to a slower function. The asymptotic version in the textbooks isn't achievable in the real world (for all but very small dimension).
There's an elementary proof using generating functions.
edit to add: If anyone's curious about it, a simplified version of the recurrence relation that's enough to exhibit this behavior (you can instantly see it if you graph this numerically) is
The curse of dimensionality times the reality that good metrics are elusive or themselves a bit cursed. Many outcomes you're engineering or product-managing toward are quite squishy, hard to define, and hard to evaluate. "Easy to use" or "can be used within 10 minutes" or "cleans up this current order form" are easy to state but hard to rate and/or hard to actionably implement as metrics.
I've built large, deep product evaluation frameworks, and it is 100% of the time a running argument with stakeholders, inside and out, "well you should have measured it this way" or "I think we should be targeting X not Y" or "why didn't you consider Z in the metric??"
The Pareto Front in practice is squishy, fuzzy, and often quite moist and moldy.
krapht 1 days ago [-]
One I spent a few months working on was pathfinding for trucks. The goal is to find dominant solutions over {shortest time, lowest cost (tolls + fuel), avg road speed variance - traffic sensitivity} and then return 3-4 routes that are equal distance from each other in this dimensional space for users to pick from.
As you say, the most useful things happen in low-dimensional spaces.
shermantanktop 1 days ago [-]
I’m sadly twitchy when I hear “Pareto” - having endured numerous middle managers suggesting they can deliver 80% of the scope in 20% of the time (unrelated to the frontier topic here). Do that at each level of an org and the nonsense multiples rapidly.
The 80/20 “rule,” as far as I know, is meant to be descriptive after the fact. It can’t be used as a planning assumption. To be fair to those managers, they don’t really mean to be rigorous. They are just trying to justify cutting scope.
CGMthrowaway 16 hours ago [-]
Using the 80/20 rule to plan, is like that other old saw "Half the money I spend on advertising is wasted. The trouble is, I don't know which half"
MarkusQ 1 days ago [-]
Managers trying to justify _cutting_ scope...
Is your planet accepting immigrants? I think I'd like it there
shermantanktop 2 hours ago [-]
They don’t cut the visible scope. They cut the monitoring, failover, automated patching, test coverage, deployment improvements, documentation, etc.
nonameiguess 24 hours ago [-]
That's the "Pareto Principle" whereas the frontier is talking about Pareto efficiency. They have the same name because both were first developed by the economist Vilfredo Pareto, but they're not actually otherwise related.
bellowsgulch 21 hours ago [-]
I don't think power law distributions are the same thing. However, actually, in different fields, power law distributions are descriptive enough you can use them as targets for abstract criteria.
> and every solution not in the set is outperformed by at least one solution in the Pareto front in every objective
Is that trying to say:
"for every solution not in the set, there exists at least one objective such that at least one solution in the Pareto set beats that solution in that objective" i.e. every non-Pareto-front solution is beaten in some objective(s) by a Pareto-front solution, however it may be unbeaten in other objectives.
Or is it:
"for every objective in the system, every solution that is not in the set is beaten in that objective by one or more Pareto-set solutions."
Or is it:
"For every solution not in the set, there exists at least one Pareto solution which beats it in every objective."
speedstyle 8 hours ago [-]
The latter. Every solution not in the set is 'dominated' (outperformed in every objective) by some specific point in the set.
kazinator 2 hours ago [-]
The way it's written is misleading. The Pareto front solutions are all dominators of the non-solutions, but not dominators of each other.
For A to dominate B, A has to be at least as good (i.e. no worse) than B in every objective under consideration and A has to be strictly better than B in at least one objective.
Every solution in the Pareto front set dominates every solution not in that set: is at least as good in all optimization parameters and strictly better in at least one.
Among the front set, there is no mutual dominance: if we pick any pair out of the set, one may be better than the other in one or more parameters, but worse in one or more. If it were not worse in one or more than the other, that other would not belong in the front set due to being dominated.
Consider a space where we have two solutions. One is no worse than the other in every objective, and strictly better in one objective. Here, our Pareto front set contains that one solution and the other one is not in the set. Yet, the one not in the set is not beaten in every objective, just in that objective where the dominator is strictly better.
cpa 1 days ago [-]
At $JOB, I use the Pareto frontier all the time.
If one option is at least as good on every relevant dimension and better on one, just pick it. That's not really a trade-off, and it shouldn't need escalation. Eg, if two SaaS tools cost the same and have similar support, but one fits your use case better, you choose that one. Otherwise, you just suck at your job!
The interesting decisions only start once you're already on the frontier, where getting more of one thing means giving up something else. If the better tool costs 50% more, now you're trading capability against cost, and that may need sign-off.
Basically, everyone should be able to get to the frontier on their own. Coordination and arbitration at higher levels of the org / between different departments should happen on the frontier, where the trade-offs involve several people or teams.
myroon5 14 hours ago [-]
(A few years outdated) AWS EC2 instance type pareto frontier:
Question: in auto racing, could one have a Pareto Front balancing single lap pace (qualifying optimization) and race pace (pace over an entire stint of e.g. 20+ laps)?
transitorykris 1 days ago [-]
You’d be looking at fuel level affecting choices of ride height, brake bias, etc. Possibly changes in line and distance travelled too. But, I’m curious how psychology can fit in here (or not. How do you measure it?). Driver concentration and confidence are important when trimming out aero or other adjustments for single lap flyers.
lorey 1 days ago [-]
Found this to display the optimal LLM choice while building evalry. It's such a useful tool, not only for thinking about it, but for visualization, too.
I'm assuming you must have discovered this through the OpenRouter LLM performance graphs.
Aachen 22 hours ago [-]
First mention in chats for me was September 2016. Not sure how that topic came up but LLMs aren't the only way to get there. It's also a rather obvious principle to come up with, at least I'm pretty sure I was doing it without thinking I should name the concept of picking the option that best fits your requirements
I think the one from hugging face is much clearer, albeit its an arena metric and a bit tricky to find the pareto view. Look for the the top Navigation Bar
(Agent Chat Code Image Video). Chat -> (dropdown) Text -> (Side panel) View as Pareto.
Nice, this is exactly what I use for the multi-objective optimizer on a quantum network simulator I'm building — scoring topologies on fidelity/latency/success rate tradeoffs.
matsemann 1 days ago [-]
My thesis many years ago was on multi-objective optimization using evolutionary algorithms (in my profile), and maintaining a wide pareto front was what most algorithms (like NSGA-II) were attempting. If all individuals cluster around a small area (in for instance a weight/strength tradeoff), you will quickly get stuck. So should select solutions to keep for further search along the whole front (for instance some solution that is very strong but unfortunately also very heavy). Maybe keep some of them as candidates even if worse (not part of the pareto front), just to keep that part of the search space alive and avoid local optima.
Of course, what's hard anyways when you have a good set of solutions that are pareto optimal, is to then choose between them. Especially as the dimensions (objectives) grow. In my example we can end up with many variants of strength/weight trade-offs that each are optimal, which one to choose?
stevefan1999 1 days ago [-]
I wonder why LLM love this word so much. Same as mint, seam, tier.
amingilani 1 days ago [-]
A seam is a place where you can alter behavior in your program without editing in
that place
“Chapter 4: The Seam Model”, Michael C. Feathers, Working Effectively with Legacy Code
pmarreck 20 hours ago [-]
A very useful concept to know!
chermi 1 days ago [-]
We used to just call that efficiency. Overusage of "pareto frontier" annoys me almost as much people talking about "electrons" instead of just saying electricity or power.
hnfwd5lqmp 22 hours ago [-]
Simple idea, big payoff
bhanu786 1 days ago [-]
may, anyone explain what is this
chriswarbo 1 days ago [-]
If we have a set of things (e.g. language models) and some measures we care about (e.g. cost, speed, whether weights are open, scores for a few benchmarks, etc.), then some of those things will be "pareto optimal" (see below) and some won't. The "pareto front" is the subset that is pareto optimal.
Some thing is "pareto optimal" when there isn't another thing that's AT LEAST AS GOOD in ALL measures, and BETTER in at least one way. For example, if we say there are no ties (for simplicity), then the cheapest language model is pareto optimal; the fastest model is pareto optimal; those which score highest on each benchmark are pareto optimal; and so on.
Tradeoffs can also be pareto optimal: for example, if the cheapest model is also slow, then there will be more pareto optimal models which are "cheapest for their speed"; and so on for other tradeoffs (e.g. fastest that achieves a certain benchmark score; cheapest model with open weights; etc.).
If you're making a decision about which thing to choose, you only need to care about those in the pareto front (since, by definition, anything that's not pareto optimal is objectively worse on at least one measure).
Pareto optimality does not compare one measure against another: something that's 10000x slower can still be pareto optimal, if it's 1% cheaper than the alternatives. To pick a "best" thing, you could give a weight/importance to each measure, and combine them into an overall score: but that's subjective, and might vary between people and tasks. In contrast, focusing on the pareto front is a way to ignore those things that will never be the best, regardless of weighting.
matsemann 1 days ago [-]
I honestly think the wikipedia article is too complicated. My own image example here as an another attempt to explain: https://imgur.com/a/5ZQIJDb
Mapping the cost of something (like an algorithm), and the time it takes (so lower is better for both). 1, 3 and 5 are all optimal in their own sense. No one is strictly better than the other, just different tradeoffs you have to choose yourself. However, you would never choose 2, because for a lower cost you could get the same result choosing 3. Same with 4, 6 and 7, they all have something that's both faster and at the same time just as cheap you could choose.
A pareto front is a bit like the classical "fast, cheap, good, choose 2". There are always tradeoffs, but if something is both slow, expensive and not better than something that's faster and cheaper, it's a bad choice, and thus not on the "pareto front".
felixguendling 1 days ago [-]
If you optimize one criterion, it's simple: lowest is best or highest is best.
If you optimize multiple criteria, all optimal trade offs between any of the selected criteria are "best" in some way.
continuational 1 days ago [-]
When you have a tradeoff between two parameters, which points dominate the others in the sense that you can't choose another point without getting less of one of the parameters.
bhanu786 1 days ago [-]
thanks, may you tell me where we can use them?
joshka 1 days ago [-]
The current thing that comes up regularly is choosing an LLM setup.
isoprophlex 1 days ago [-]
"what's the family of optimal choices when you have multiple dimensions to rank on?"
Say a race vehicle has acceleration, top speed as defining parameters. Some are slow but accelerate hard, others need a long time to reach very high top speeds. Others are in between, or just flat out bad at both.
The pareto frontier is the set of vehicles that are best: pick one from the frontier and you can be sure that for it's given top speed, none accelerate faster. And vice versa, pick one with a given acceletation and you are sure none have a better top speed
ChrisMarshallNY 1 days ago [-]
Basically, prioritization.
It’s really that simple.
Eschew obfuscation.
Lerc 1 days ago [-]
Almost the complete opposite of prioritisation.
The Pareto points are where you sacrifice the least of anything to get the most of everything.
There's the saying about buying computers. Good, Cheap, Fast, pick any two. That's where you would prioritise.
If someone makes something that better, cheaper, and faster, or even pretty close to the best on two of those and clearly better on the other. It's a Pareto point.
Over time computers are getting better, cheaper and faster (software notwithstanding). The leading edge of that advance of all of the things is the Pareto front.
jrrv 1 days ago [-]
Given this in the TFA
> a Pareto front represents the set of solutions where no solution outperforms any other solution in the set at every objective
I do not believe you are correct when you say
> something that better, cheaper, and faster, or even pretty close to the best on two of those and clearly better on the other. It's a Pareto point.
Since that would outperform on every objective
GP's point that it's prioritisation does not seem incorrect to me. Prioritisation involves considering trade-offs of various approaches and deciding which aspects & attributes to optimise for, at the expense of others.
Lerc 1 days ago [-]
[dead]
ChrisMarshallNY 1 days ago [-]
I’m not sure how that’s the opposite of prioritizing, but if I’m wrong, then I’ll happily admit it.
We choose our items/workflows/technologies/whatever, so we get the best/most efficient/most effective/whatever, across the widest possible set.
Sounds like prioritizing, to me, but I’m just a dumb hick, so I suppose I can be wrong.
Lerc 1 days ago [-]
Prioritising is when you choose something over another. A drag racer prioritises time to travel a quarter mile.
Going for the Pareto is when you elect not to prioritise. It is explicitly deciding to not choose one property over another ant to keep everything as much as you can.
ChrisMarshallNY 22 hours ago [-]
I read it as finding the point of maximum effectiveness. The point at which the most is done for the most.
Getting to that point can be calculated (in some cases), but I suspect most folks get there by trial and error. Finding out what is effective, and what is not, and choosing what is effective, over what is not, until there's no longer a choice. That often becomes tribal knowledge, and is handed down. There's always someone trying to improve it, and when they figure it out, that gets added to the tribal knowledge. Basically, that's how nature does it, so there's some serious prior art. Natural Selection is brutal prioritization.
In Morocco, they used to announce the end of the Ramadan fast, by holding up a black thread and a white thread, and waiting until they could not tell the difference.
Then, they'd fire a cannon, and everybody would dig into some awesome soup. Sort of the same thing.
stevefan1999 1 days ago [-]
No, it's more about maximizing the utility function and finding the "shield" where you'd start getting diminishing return beyond that
bhanu786 1 days ago [-]
you have confused me
voidhorse 1 days ago [-]
One nuance that people sometimes miss is that pareto optimality in the continuous case and discrete case are distinct. Using continuous case algorithms on discrete feasible set optimization problems will make you miss the interior optimal points--only extremal/supported points on the positive orthant hull are identified by the continuous algos.
Matthias Ehrgott's books on multicriteria optimization explain Pareto efficiency very well without sacrificing rigor. I think they do a better job than this article.
solomonb 23 hours ago [-]
Now I want a hat with the Agnostic Front logo but that says Pareto Front.
dancemethis 12 hours ago [-]
The real Pareto is the lawyer who was the victim of the Telerj prank call in the 80s. Don't be fooled!
I use conditional formatting to color cells according to the probability that I can lift them—if I lifted 50kg for 10 reps then I can definitely do 50kg for 9 reps, so that cell is green. But if e1RM(50,10) > e1RM(40,15) then I can probably do that too so it's light green. The visualization naturally becomes Pareto-like.
If I'm feeling strong I can aim for higher weight, lower reps. Or if I'm feeling weak I can close out a (weight, reps) that's below my current e1RM but I haven't accomplished yet. The end result is that I'm always "accomplishing" some sort of PR no matter how I feel.
I call this e1RM Bingo.
My main finding for “pick whatever weight you want today” was that picking a lot of different weights made the curve less identifiable, so my latest iteration encourages you to pick a ladder for a few sentinel exercises per mesocycle in order to improve the statistical power. In addition, strength improves more quickly at >80% of 1RM, and hypertrophy depends on proximity to failure, so if you pick a lower weight, you really need to go to failure, which burns you out for the rest of your session, where leaving 1-2 reps in reserve is probably sufficient for hypertrophy and leaves a lot more gas in the tank for the rest of the session. Definitely open to suggestion/discussion here.
https://curvefit.app (it runs on Cloudflare free tier, so I won’t have to start running ads or charging until I hit a couple thousand users)
I'm not one to believe in all the one-shot hype, but this was pretty good.
If pursued, good luck!
Please don't make an app based on this.
Some nuance here: the latest research shows that proximity to failure is the main hypertrophy driver regardless of load and rep count; high rep count makes proximity to failure harder to gauge; so high load/low reps close to failure is probably better for hypertrophy (there are other good reasons to do higher reps/lower load work though)
If you want the most 'optimal' form of this (aka, hell on earth), you should purchase a rowing machine. Being able to engage with very aggressive, full-body exercise every single day without exceptions is almost like cheating biology. You can maintain a 2-3x VO2 max premium over your peers with very little risk of injury.
But they go even a step further, they extend into 3 dimensions to also add body weight as a variable. So your graph would really have to be a 3D volume. Because different levels of body weight have different capabilities.
Maybe in vein but did anyone already figure this one out? The closest I got was PT sans, open-licensed commissioned by the Russian ministry for communication (I found it surprising that a country that doesn't use Latin script made the best font!), but it's not widely shipped so you need to figure out how to include font files whenever you want to use it
"The Pareto Front today claimed responsiblity for...."
[1] - http://montypython.50webs.com/scripts/Life_of_Brian/8.htm
Anyways I'll namedrop Iosevka as perfect monospace font for working on 13" laptop
https://rsms.me/inter/
ss02 disambiguation seems to be the one I'd be wanting to turn on, with tnum for monospace numbers being a good option as well that I hadn't even realised I wanted from a font!
Tabular numbers are awesome!
They apparently released Hyperlegible Next in 2025 which, flipping between tabs on Google Fonts (since the original website doesn't show the fonts), is nearly identical but has five new weight settings (nobody should imo ever use thin fonts though, it noticeably harms readability for me and my sight is only the tiniest bit below normal vision, but ok it's an option) and improved kerning (the original font had extremely little space between 'll', for example)
The 2025 version sadly doesn't ship with my version of TexLive, but the original (from 2020) already does so that makes it easy to use as well! Cool stuff, thanks for the tip :)
When explaining it to some coworkers, I stumbled on a fairly intuitive explanation: "I've run farther before, and I've run faster before, but I've never run _this_ far, _this fast."
There was some pushback about why not just call it a PR (personal record), but I would only use that term for fixed distances (1mi, 5k, 10k, etc.) or a consistent route that I've run many times before. Nobody would say "I set my 7.40 mile PR today." More importantly, it misses the comparison to all farther (and faster) runs—it's not exciting to set a 5k PR just because you've barely run that distance before, and the pace is actually slower that a 10k you've done.
(Had a Pareto run of 7.40 miles @ 6:28/mi last week!)
A point is Pareto so long as it is non-dominated—that is, you're not looking for dominating points, you're looking for points that are "no worse" than all others, in all criteria, when you consider that point as a reference.
So your Pareto Runs are indeed Pareto points. However, your run with your fastest possible speed, even if your distance was really bad, is also still a Pareto efficient point.
(Taking >= as more efficient here) By definition, the point A is Pareto if there is no point B such that in all criteria, B >= A, and for at least one criteria B > A. Take the run with the best speed. It is Pareto because we cannot find a single point B that satisfies both of these conditions. Your "Pareto Run" doesn't satisfy this set of conditions because it is worse in terms of speed, even if it has better distance than the max speed point.
The only way your Pareto runs would be the only Pareto points in your record is if they simultaneously hit maxima for distance and speed when compared to all other points. So, for them to be the sole Pareto point, the clause ""I've run farther before, and I've run faster before..." would have to be false! The point would have to break both your all time records to be the solitary Pareto point. With running, because of how speed and distance are related this will basically never happen.
The definition of Pareto efficiency is essentially negative in nature--it's not about finding specific dominating points, it's about finding points that are not dominated by any others on any criterion, period. All criteria are weighted equally in the search for Pareto points. It doesn't build in any weighting like considering maximum across criteria as "better" than points that only maximize one criteria. For a "biobjective" problem like your runs, the Pareto set will always contain the points (MAX, -) and (-, MAX)--they may not be unique over the criteria but there will always be at least one representative for each, I believe.
ChatGPT 5.6 Luna on the right (cheaper) cover most of the frontier, with a point for Deepseek flash, and higher performance overlapping heavily between 5.6 Sol and Fable.
That DeepSeek point will probably move back towards Luna as deepseek announced a "significant" price increase coming to their API [1], which kind of demonstrates that beating the Pareto frontier is where the difficulty actually is).
[1] https://www.bloomberg.com/news/articles/2026-08-06/deepseek-...
I think it's great and hope the price can stay the same.
Going from 81GB of weights to 79GB of weights can mean a 50% reduction in GPU capacity required.
If you can fit a model in just one GPU (or rack) as opposed to across an entire datacenter, the latency gains can be substantial too. If you can reduce token latency by half, that would double the amount of customers you could support.
The objectives are matching arguments to parameters.
A set of functions is identified among the candidates: those that are possible for the call at all, like having a compatible number of parameters.
Essentially, the overload rule says that the Pareto front set of candidates must contain one member, otherwise the call is considered ambiguous, and diagnosable rule violation.
The objectives being optimized are individual parameter positions, each in the dimension of suitability: being a better match.
One candidate is better than another if it is no worse a type match in every parameter, and strictly better in at least one parameter.
As the number of objectives (dimensions) increases, the number of samples you need to cover the frontier increases exponentially. You will very rarely find solutions that actually dominate other solutions in many practical optimization scenarios. With 2 dimensions you have a 25% chance of domination. With 10 dimensions it's a .098% chance.
The most useful cases I've seen tend to occur where we just optimize for two things at once. The chances of domination are high, it's easy to visualize and very efficient to implement. As we get into higher dimensional spaces, things get weird really fast.
The geometric problem of computing a d-dimensional Pareto set of cardinality n
https://en.wikipedia.org/wiki/Maxima_of_a_point_set
has a truly weird property not covered by the computational complexity discussion on that page. It says there's an algorithm achieving O(n log(n)^(d-3) log log n), which is true and also a lie. The algorithm that achieves that asymptotic form is a galactic algorithm; and not an ordinary one in the sense of "has a large constant multiplicative factor", but one with this property (I've never found any other algorithm which exhibits it):
The runtime is within a bounded constant factor of n^2, for all n up to some critical N whose size is exponential in d (I think it was exactly 2^d or something).
I.e. the runtime has "two shapes": it's purely quadratic up to a galactically-large constant, and thereafter has a transition into to a slower function. The asymptotic version in the textbooks isn't achievable in the real world (for all but very small dimension).
There's an elementary proof using generating functions.
edit to add: If anyone's curious about it, a simplified version of the recurrence relation that's enough to exhibit this behavior (you can instantly see it if you graph this numerically) is
I've built large, deep product evaluation frameworks, and it is 100% of the time a running argument with stakeholders, inside and out, "well you should have measured it this way" or "I think we should be targeting X not Y" or "why didn't you consider Z in the metric??"
The Pareto Front in practice is squishy, fuzzy, and often quite moist and moldy.
As you say, the most useful things happen in low-dimensional spaces.
The 80/20 “rule,” as far as I know, is meant to be descriptive after the fact. It can’t be used as a planning assumption. To be fair to those managers, they don’t really mean to be rigorous. They are just trying to justify cutting scope.
Is your planet accepting immigrants? I think I'd like it there
Is that trying to say:
"for every solution not in the set, there exists at least one objective such that at least one solution in the Pareto set beats that solution in that objective" i.e. every non-Pareto-front solution is beaten in some objective(s) by a Pareto-front solution, however it may be unbeaten in other objectives.
Or is it:
"for every objective in the system, every solution that is not in the set is beaten in that objective by one or more Pareto-set solutions."
Or is it:
"For every solution not in the set, there exists at least one Pareto solution which beats it in every objective."
For A to dominate B, A has to be at least as good (i.e. no worse) than B in every objective under consideration and A has to be strictly better than B in at least one objective.
Every solution in the Pareto front set dominates every solution not in that set: is at least as good in all optimization parameters and strictly better in at least one.
Among the front set, there is no mutual dominance: if we pick any pair out of the set, one may be better than the other in one or more parameters, but worse in one or more. If it were not worse in one or more than the other, that other would not belong in the front set due to being dominated.
Consider a space where we have two solutions. One is no worse than the other in every objective, and strictly better in one objective. Here, our Pareto front set contains that one solution and the other one is not in the set. Yet, the one not in the set is not beaten in every objective, just in that objective where the dominator is strictly better.
If one option is at least as good on every relevant dimension and better on one, just pick it. That's not really a trade-off, and it shouldn't need escalation. Eg, if two SaaS tools cost the same and have similar support, but one fits your use case better, you choose that one. Otherwise, you just suck at your job!
The interesting decisions only start once you're already on the frontier, where getting more of one thing means giving up something else. If the better tool costs 50% more, now you're trading capability against cost, and that may need sign-off.
Basically, everyone should be able to get to the frontier on their own. Coordination and arbitration at higher levels of the org / between different departments should happen on the frontier, where the trade-offs involve several people or teams.
https://github.com/PatMyron/cloud#compute--memory-unit-price...
Example: Which LLM gives me the best ELI5 explanations for a given price. https://evalry.com/benchmarks/explain-like-i-m-5-321
https://artificialanalysis.ai/#intelligence-comparison-tabs
https://huggingface.co/spaces/lmarena-ai/arena-leaderboard
Of course, what's hard anyways when you have a good set of solutions that are pareto optimal, is to then choose between them. Especially as the dimensions (objectives) grow. In my example we can end up with many variants of strength/weight trade-offs that each are optimal, which one to choose?
“Chapter 4: The Seam Model”, Michael C. Feathers, Working Effectively with Legacy Code
Some thing is "pareto optimal" when there isn't another thing that's AT LEAST AS GOOD in ALL measures, and BETTER in at least one way. For example, if we say there are no ties (for simplicity), then the cheapest language model is pareto optimal; the fastest model is pareto optimal; those which score highest on each benchmark are pareto optimal; and so on.
Tradeoffs can also be pareto optimal: for example, if the cheapest model is also slow, then there will be more pareto optimal models which are "cheapest for their speed"; and so on for other tradeoffs (e.g. fastest that achieves a certain benchmark score; cheapest model with open weights; etc.).
If you're making a decision about which thing to choose, you only need to care about those in the pareto front (since, by definition, anything that's not pareto optimal is objectively worse on at least one measure).
Pareto optimality does not compare one measure against another: something that's 10000x slower can still be pareto optimal, if it's 1% cheaper than the alternatives. To pick a "best" thing, you could give a weight/importance to each measure, and combine them into an overall score: but that's subjective, and might vary between people and tasks. In contrast, focusing on the pareto front is a way to ignore those things that will never be the best, regardless of weighting.
Mapping the cost of something (like an algorithm), and the time it takes (so lower is better for both). 1, 3 and 5 are all optimal in their own sense. No one is strictly better than the other, just different tradeoffs you have to choose yourself. However, you would never choose 2, because for a lower cost you could get the same result choosing 3. Same with 4, 6 and 7, they all have something that's both faster and at the same time just as cheap you could choose.
A pareto front is a bit like the classical "fast, cheap, good, choose 2". There are always tradeoffs, but if something is both slow, expensive and not better than something that's faster and cheaper, it's a bad choice, and thus not on the "pareto front".
Say a race vehicle has acceleration, top speed as defining parameters. Some are slow but accelerate hard, others need a long time to reach very high top speeds. Others are in between, or just flat out bad at both.
The pareto frontier is the set of vehicles that are best: pick one from the frontier and you can be sure that for it's given top speed, none accelerate faster. And vice versa, pick one with a given acceletation and you are sure none have a better top speed
It’s really that simple.
Eschew obfuscation.
The Pareto points are where you sacrifice the least of anything to get the most of everything.
There's the saying about buying computers. Good, Cheap, Fast, pick any two. That's where you would prioritise.
If someone makes something that better, cheaper, and faster, or even pretty close to the best on two of those and clearly better on the other. It's a Pareto point.
Over time computers are getting better, cheaper and faster (software notwithstanding). The leading edge of that advance of all of the things is the Pareto front.
> a Pareto front represents the set of solutions where no solution outperforms any other solution in the set at every objective
I do not believe you are correct when you say
> something that better, cheaper, and faster, or even pretty close to the best on two of those and clearly better on the other. It's a Pareto point.
Since that would outperform on every objective
GP's point that it's prioritisation does not seem incorrect to me. Prioritisation involves considering trade-offs of various approaches and deciding which aspects & attributes to optimise for, at the expense of others.
We choose our items/workflows/technologies/whatever, so we get the best/most efficient/most effective/whatever, across the widest possible set.
Sounds like prioritizing, to me, but I’m just a dumb hick, so I suppose I can be wrong.
Going for the Pareto is when you elect not to prioritise. It is explicitly deciding to not choose one property over another ant to keep everything as much as you can.
Getting to that point can be calculated (in some cases), but I suspect most folks get there by trial and error. Finding out what is effective, and what is not, and choosing what is effective, over what is not, until there's no longer a choice. That often becomes tribal knowledge, and is handed down. There's always someone trying to improve it, and when they figure it out, that gets added to the tribal knowledge. Basically, that's how nature does it, so there's some serious prior art. Natural Selection is brutal prioritization.
In Morocco, they used to announce the end of the Ramadan fast, by holding up a black thread and a white thread, and waiting until they could not tell the difference.
Then, they'd fire a cannon, and everybody would dig into some awesome soup. Sort of the same thing.
Matthias Ehrgott's books on multicriteria optimization explain Pareto efficiency very well without sacrificing rigor. I think they do a better job than this article.