Repository navigation
[FileFormats.MPS] slowing reading of .mps file #2888
Description
Activity
- changed the title
[-]Slowing reading of .mps file[/-][+][FileFormats.MPS} slowing reading of .mps file[/+]on Dec 2, 2025 There are known overheads associated with
read_from_file. See my recent talk on MOI: https://youtu.be/M31xoZGyj9w?si=Ma60aWbEAV4RAePt&t=1667We could perhaps be a bit faster, but the times aren't really directly comparable between JuMP and HiGHS because JuMP is doing significantly more work.
A better question to ask is: why do you care that reading this file takes a few seconds? In most cases this shouldn't be a significant runtime part of any optimization-related workflows.
- changed the title
[-][FileFormats.MPS} slowing reading of .mps file[/-][+][FileFormats.MPS] slowing reading of .mps file[/+]on Dec 2, 2025 Using
MOI.readModel()is indeed faster thanJuMP.read_from_file(). As shown in the following:
julia> @time dest = MOI.FileFormats.Model(format = MOI.FileFormats.FORMAT_MPS);julia> @time MOI.read_from_file(dest, "2025-1126-SCUC-noncon.mps");
22.695831 seconds (150.83 M allocations: 9.657 GiB, 29.85% gc time)A better question to ask is: why do you care that reading this file takes a few seconds? In most cases this shouldn't be a significant runtime part of any optimization-related workflows.
That is because in this scenario, the reading of mps files is used in a self-defined primal heuristic that aims to find a good feasible solution as soon as possible, instead of building a model, passing to solver and waiting for the result. Moreover, this heuristic could be revoked multiple times in one solution procedure. Therefore, a 20-second improvement in the speed of reading from .mps file every time is not trivial.
I think that MOI could offer unique values in this type of usage, where both computational speed and flexibility of manipulating the model are desired. Tools like Pyomo are very flexible but could be slow, while the C API of HiGHS is fast but it is not easy to manipulate the model with it. By manipulating the model, I mean adding new binary variables to model as well as new constraints bonding the new variables and existing binary variables.
That is because in this scenario, the reading of mps files is used in a self-defined primal heuristic
Why is reading an MPS file part of the heuristic?
Why is reading an MPS file part of the heuristic?
Two reasons:
- use another solver to solve the sub-MIP. For example, modify the sub-MIP based heuritisic in HiGHS so that SCIP is used to solve the sub-MIP problem
- easier to change the model substaintially. For example, it seems to me that the following requires quite some work to implement with the C API of HiGHS:
- Step 1: add some binary variables to the model.
- Step 2: add constraints, each of which binds one new variable with an existing variable.
- Step 3: add one constraint requiring that the new variables sum to 1.
It seems to me it is a lot easier to implemt this in MOI/JuMP than HiGHS.jl (using the C API), which makes the code easier to maintain and less prone to bugs.
I have so many questions.
You're editing the sub-MIP heuristic in HiGHS by changing the C++ code, and as part of that, you're writing the MPS file to disk, starting Julia, reading the MPS file, doing a bunch of stuff, then writing back a file of solutions, which you then load back into C++ and keep going?
Is the purpose just research to help prototype new heuristic algorithms? It can't be performance critical, right?
Would it be sufficient to have access to the heuristic callback for HiGHS?
We just generally haven't prioritised file I/O because you should avoid it at all costs.
Can you share the MPS file?
I made a small improvement in #2892
Reacted by JunchengLiamLiI have so many questions.
You're editing the sub-MIP heuristic in HiGHS by changing the C++ code, and as part of that, you're writing the MPS file to disk, starting Julia, reading the MPS file, doing a bunch of stuff, then writing back a file of solutions, which you then load back into C++ and keep going?
Is the purpose just research to help prototype new heuristic algorithms? It can't be performance critical, right?
Would it be sufficient to have access to the heuristic callback for HiGHS?
We just generally haven't prioritised file I/O because you should avoid it at all costs.
Oh, I just realised that I probably did not explain it well.
I am actually using the heuristic callback for HiGHS, with its C API called from Julia. I completely agree with you that it does not make sense to change the C++ code and call Julia at the same time.
The reasons why I am using files I/O at the moment:
- I am struggling with building model from the C function Highs_passLp() and thus using the naive alternative of writing and reading model from the disk. I will certainly try to make Highs_passLp() work and replace the file I/O method.
- I am modifying the model heavily at some point of the heuristic algorithm, which is seemingly quite hard to implement. That is why I resort to MOI/JuMP for that particular step. I do not know how to build a MOI model from a highs instance in C, apart from the file I/O approach.
Can you share the MPS file?
No, the file is confidential. However, I can share a mps file with similar size, created from UnitCommitment.jl
Is this really just asking for the JuMP-level heuristic callback support to be added to HiGHS.jl? jump-dev/HiGHS.jl#313
Well, this is a different one. Just imagine that if I want to use cuPDLPx https://github.com/MIT-Lu-Lab/cuPDLPx inside a heuristic callback of HiGHS. Just thought using the mps file would be a convenient way to communicate between cuPDLPx and HiGHS. You might ask why not pass the original mps file to cuPDLPx. That is because I want cuPDLPx to solve a problem that has already been reduced by some components in HiGHS (e.g., MIP presolving). Just wondering if that makes sense.
I'm still confused what your goal/motivation is. Is this for a specific model, or some generic heuristic algorithm? The heuristic callbacks for a solver aren't called for sub-mips, so you only ever work in the original space of the model. They give you a fractional solution and ask for an integer feasible solution in return. How does cuPDLPx fit into this scenario?
If you want to use specific parts of HiGHS like the presolve via the C API, then we'd generally suggest that you query the model in memory via the C AP (like I see you're doing in the other issue 👍). You might also consider investigating something like PaPILO.
Reacted by JunchengLiamLiI agree with you. I am now able to query the model in memory via the C API. This issue can be also left for the future.
I'm still confused what your goal/motivation is.
As for the motivation, it could be quite useful for solving MILP problems. Let me explain using the following example. Imagine a large MILP model, whose size is reduced by a half in the presolving process. If the LP relaxation of the reduced model is solved to a certain feasilbility tolerance (e.g. 1e-4) by cuPDLPx quickly, which gives a lower bound (the original model is a minimisation), and then running a primal heuristic using the LP relaxation solution, which gives an upper bound. With both a lower and an upper bound, reduced-cost fixing can be performed on the model to further reduce it, which results in a stronger model in the subsequent branch-and-bound tree. Just wondering if this makes sense.
Thanks very much for the conversion! I learned quite a few things from it.
I understand what you're trying to do. But it seems more of a solver engineering task. You're preserving, calling solvers, running heuristics, etc. It isn't obvious that JuMP should be part of that, or that reading an MPS file is necessary or on the hot-path so that reading time makes a difference!
I don't disagree that we can make the MPS reader faster. But it requires spending a day or two of engineering time, which is hard to justify when it is used infrequently. There's also the inherent design issue that we first read into
MPS.Modeland then copy to the user, so it's never going to be as performant asHighs_readModel.I'll add the "help wanted" flag to this issue. If anyone reading wants to have a go at optimizing the code, have at it! I'll happily review a pull request with benchmarks.
Reacted by JunchengLiamLi- addedStatus: help wantedThis issue needs could use help from the communityThis issue needs could use help from the communitySubmodule: FileFormatsAbout the FileFormats submoduleAbout the FileFormats submodule
on Dec 9, 2025
The
read_from_file()function is quite slow on my .mps file. It takes 33 seconds to build a model from the .mps file, while the C_API of HiGHS (call from Julia) takes only about the 8 seconds. Please see the below:@time model_JuMP = read_from_file("2025-1126-SCUC-noncon.mps");33.343285 seconds (236.12 M allocations: 14.443 GiB, 25.32% gc time)@time HiGHS.Highs_readModel(model_HiGHS,"2025-1126-SCUC-noncon.mps");7.832455 seconds (1.50 k allocations: 78.234 KiB, 0.05% compilation time)Just wondering if there is any way to build a JuMP.Model faster from .mps files.
Many thanks!