Yes. -gc branch should be updated but I personal want to wait for the v3.19.2 stable version to be out then rebase on it, so it is late.
And while the -new branch once appeared in git which I used to sync up git tree between desktop and notebook make some trouble for friends who have tried it. This remind me to publish codes carefully.
There is no new commits for -gc branch, just minor bug/comment fixes. I have rewritten some commit titles and tag with [Fix]/[Sync], so it would be easier to tell what the commit is used for.
And, I have done one *make no more friend* thing, I have delete the origin linx-3.19.y-gc branch and create a new one, with the same name. Because I don't want to keep the old merge commits info in the branch. Any one know how to make it clean please let me know.
So, you may have trouble to pull out this new linx-3.19.y-gc branch, because your local copy is different and no more existed in "server" side. Simply delete the old remote linx-3.19.y-gc branch in your git tree and pull again, then you will be fine.
Branch url:
https://bitbucket.org/alfredchen/linux-gc/commits/branch/linux-3.19.y-gc
Monday, March 23, 2015
Monday, February 9, 2015
VRQ: About Issue found in 0.3
Issue
Thanks Manuel for testing -vrq and report an issue that while background(SCHED_BATCH, nice 19) workload is running, one of the cpu failed to pick up normal/system tasks.
The detail of the issue can be found from http://cchalpha.blogspot.com/2014/12/vrq-03-updates.html
Cause
After asking Manuel about his usage for a few rounds, I finally be able to reproduced the issue and used bisect to find out that "bfs: vrq: RQ niffy solution." which introduce the issue.
The root cause is that the difference of niffy of each cpu goes very large after system keep on running, I recorded 6+ seconds difference after system up for 16+hours in my system. This cause the tasks which run on the cpu with lower niffy has earlier deadline than others, so it failed to pick up other normal/system load.
Solution
Simplely revert the commit can fix the issue but it is against the intention of commit "bfs: vrq: RQ niffy solution.", to make update_clocks() grq lock free.
After testing, I found that rq->clock which based on sched_clock_cpu() is stable on my system and the difference among cpus are small enough for deadline calculation. So I give the rq->clock a try and remove sched_clock sanity checks. I also notice that CK add the sanity check for "crazy sched_clock interface", so there may be some unexpected behaviors on some hardware, specially old machines, if this unexpected behaviors is still popular, I will add some kind of sanity check back.
The new solution will be posted with the incoming 3.19 -vrq branch, as this bug is found on -vrq and I would like the -vrq branch stay on itself for one more release before merging them into -gc branch.
Thanks Manuel for testing -vrq and report an issue that while background(SCHED_BATCH, nice 19) workload is running, one of the cpu failed to pick up normal/system tasks.
The detail of the issue can be found from http://cchalpha.blogspot.com/2014/12/vrq-03-updates.html
Cause
After asking Manuel about his usage for a few rounds, I finally be able to reproduced the issue and used bisect to find out that "bfs: vrq: RQ niffy solution." which introduce the issue.
The root cause is that the difference of niffy of each cpu goes very large after system keep on running, I recorded 6+ seconds difference after system up for 16+hours in my system. This cause the tasks which run on the cpu with lower niffy has earlier deadline than others, so it failed to pick up other normal/system load.
Solution
Simplely revert the commit can fix the issue but it is against the intention of commit "bfs: vrq: RQ niffy solution.", to make update_clocks() grq lock free.
After testing, I found that rq->clock which based on sched_clock_cpu() is stable on my system and the difference among cpus are small enough for deadline calculation. So I give the rq->clock a try and remove sched_clock sanity checks. I also notice that CK add the sanity check for "crazy sched_clock interface", so there may be some unexpected behaviors on some hardware, specially old machines, if this unexpected behaviors is still popular, I will add some kind of sanity check back.
The new solution will be posted with the incoming 3.19 -vrq branch, as this bug is found on -vrq and I would like the -vrq branch stay on itself for one more release before merging them into -gc branch.
Wednesday, December 31, 2014
VRQ 0.3 updates
In the last day of 2014, I will like to announce the 0.3 updates of VRQ solution for BFS.
It's almost a rework to address the lock dependency issues which caused by introduction of rq lock in VRQ. I'll like to do more clean up during 3.18 release, if it goes well, it will be moved to -gc branch.
Here are the test result of CFS, BFS original, Baseline(the VRQ based on) and the VRQ. It shows that after three release of VRQ, it can be better than original BFS in all kinds of workload and comparable against CFS(again? As I remember correctly, BFS is better than CFS in compiling test some release ago). But most important of all, VRQ open the opportunity to further improvements, I will give these new ideas a try next year.
Happy new Year 2015.
#1 50% task/cores ratio workload test
3.18_CFS_50
5m18.385s
5m17.836s
5m17.783s
5m17.765s
5m17.730s
5m17.663s
5m17.596s
5m17.566s
5m17.487s
5m17.455s
3.18_0460_50
5m20.448s
5m20.384s
5m20.374s
5m20.286s
5m20.245s
5m20.192s
5m20.166s
5m20.155s
5m20.150s
5m20.093s
3.18_Baseline_50
5m17.975s
5m17.552s
5m17.512s
5m17.492s
5m17.485s
5m17.482s
5m17.480s
5m17.475s
5m17.457s
5m17.342s
3.18_VRQ_50
5m17.350s
5m17.339s
5m17.318s
5m17.297s
5m17.284s
5m17.282s
5m17.248s
5m17.220s
5m17.209s
5m17.145s
#2 100% task/cores ratio workload test
3.18_CFS_100
2m50.047s
2m50.037s
2m49.825s
2m49.750s
2m49.749s
2m49.744s
2m49.706s
2m49.703s
2m49.673s
2m49.633s
3.18_0460_100
2m51.933s
2m50.640s
2m50.431s
2m50.424s
2m50.421s
2m50.386s
2m50.362s
2m50.267s
2m50.248s
2m50.129s
3.18_Baseline_100
2m51.862s
2m50.527s
2m50.506s
2m50.400s
2m50.370s
2m50.361s
2m50.283s
2m50.282s
2m50.189s
2m50.146s
3.18_VRQ_100
2m50.944s
2m49.812s
2m49.797s
2m49.700s
2m49.683s
2m49.672s
2m49.653s
2m49.649s
2m49.613s
2m49.506s
#3 150% task/cores ratio workload test
3.18_CFS_150
2m53.382s
2m53.366s
2m53.328s
2m53.326s
2m53.326s
2m53.310s
2m53.307s
2m53.262s
2m53.208s
2m53.127s
3.18_0460_150
2m57.280s
2m56.710s
2m56.124s
2m55.860s
2m55.843s
2m55.725s
2m55.646s
2m55.643s
2m55.597s
2m55.582s
3.18_Baseline_150
2m57.100s
2m55.907s
2m55.796s
2m55.788s
2m55.755s
2m55.749s
2m55.740s
2m55.736s
2m55.732s
2m55.726s
3.18_VRQ_150
2m55.449s
2m53.168s
2m52.112s
2m51.898s
2m51.649s
2m51.539s
2m51.527s
2m51.371s
2m51.272s
2m51.270s
#4 100%+50%(IDLE) task/cores ratio workload tests
3.18_CFS_100_50
2m55.069s
2m53.714s
2m53.638s
2m53.586s
2m53.474s
2m53.466s
2m53.313s
2m53.280s
2m53.196s
2m53.187s
3.18_0460_100_50
2m50.730s
2m50.713s
2m50.651s
2m50.620s
2m50.615s
2m50.598s
2m50.560s
2m50.548s
2m50.460s
2m50.457s
3.18_Baseline_100_50
2m50.826s
2m50.691s
2m50.683s
2m50.632s
2m50.601s
2m50.574s
2m50.555s
2m50.549s
2m50.549s
2m50.507s
3.18_VRQ_100_50
2m49.822s
2m49.799s
2m49.766s
2m49.706s
2m49.673s
2m49.631s
2m49.623s
2m49.587s
2m49.583s
2m49.564s
It's almost a rework to address the lock dependency issues which caused by introduction of rq lock in VRQ. I'll like to do more clean up during 3.18 release, if it goes well, it will be moved to -gc branch.
Here are the test result of CFS, BFS original, Baseline(the VRQ based on) and the VRQ. It shows that after three release of VRQ, it can be better than original BFS in all kinds of workload and comparable against CFS(again? As I remember correctly, BFS is better than CFS in compiling test some release ago). But most important of all, VRQ open the opportunity to further improvements, I will give these new ideas a try next year.
Happy new Year 2015.
#1 50% task/cores ratio workload test
3.18_CFS_50
5m18.385s
5m17.836s
5m17.783s
5m17.765s
5m17.730s
5m17.663s
5m17.596s
5m17.566s
5m17.487s
5m17.455s
3.18_0460_50
5m20.448s
5m20.384s
5m20.374s
5m20.286s
5m20.245s
5m20.192s
5m20.166s
5m20.155s
5m20.150s
5m20.093s
3.18_Baseline_50
5m17.975s
5m17.552s
5m17.512s
5m17.492s
5m17.485s
5m17.482s
5m17.480s
5m17.475s
5m17.457s
5m17.342s
3.18_VRQ_50
5m17.350s
5m17.339s
5m17.318s
5m17.297s
5m17.284s
5m17.282s
5m17.248s
5m17.220s
5m17.209s
5m17.145s
#2 100% task/cores ratio workload test
3.18_CFS_100
2m50.047s
2m50.037s
2m49.825s
2m49.750s
2m49.749s
2m49.744s
2m49.706s
2m49.703s
2m49.673s
2m49.633s
3.18_0460_100
2m51.933s
2m50.640s
2m50.431s
2m50.424s
2m50.421s
2m50.386s
2m50.362s
2m50.267s
2m50.248s
2m50.129s
3.18_Baseline_100
2m51.862s
2m50.527s
2m50.506s
2m50.400s
2m50.370s
2m50.361s
2m50.283s
2m50.282s
2m50.189s
2m50.146s
3.18_VRQ_100
2m50.944s
2m49.812s
2m49.797s
2m49.700s
2m49.683s
2m49.672s
2m49.653s
2m49.649s
2m49.613s
2m49.506s
#3 150% task/cores ratio workload test
3.18_CFS_150
2m53.382s
2m53.366s
2m53.328s
2m53.326s
2m53.326s
2m53.310s
2m53.307s
2m53.262s
2m53.208s
2m53.127s
3.18_0460_150
2m57.280s
2m56.710s
2m56.124s
2m55.860s
2m55.843s
2m55.725s
2m55.646s
2m55.643s
2m55.597s
2m55.582s
3.18_Baseline_150
2m57.100s
2m55.907s
2m55.796s
2m55.788s
2m55.755s
2m55.749s
2m55.740s
2m55.736s
2m55.732s
2m55.726s
3.18_VRQ_150
2m55.449s
2m53.168s
2m52.112s
2m51.898s
2m51.649s
2m51.539s
2m51.527s
2m51.371s
2m51.272s
2m51.270s
#4 100%+50%(IDLE) task/cores ratio workload tests
3.18_CFS_100_50
2m55.069s
2m53.714s
2m53.638s
2m53.586s
2m53.474s
2m53.466s
2m53.313s
2m53.280s
2m53.196s
2m53.187s
3.18_0460_100_50
2m50.730s
2m50.713s
2m50.651s
2m50.620s
2m50.615s
2m50.598s
2m50.560s
2m50.548s
2m50.460s
2m50.457s
3.18_Baseline_100_50
2m50.826s
2m50.691s
2m50.683s
2m50.632s
2m50.601s
2m50.574s
2m50.555s
2m50.549s
2m50.549s
2m50.507s
3.18_VRQ_100_50
2m49.822s
2m49.799s
2m49.766s
2m49.706s
2m49.673s
2m49.631s
2m49.623s
2m49.587s
2m49.583s
2m49.564s
Wednesday, November 26, 2014
-gc updates and VRQ benchmark in 3.17
There is BFS 0458 release for 3.17, I have synced up -gc with it then rebase -gc branch with v3.17.3 then v3.17.4, you can fetch them from my git.
As 0458 BFS released, I kicked off another new branch of testing to get the result of BFS 0458, Baseline of VRQ patch and the VRQ patch.
The result highlighted are
a. For low workload, Baseline&VRQ is ~3 seconds(1%) faster than original BFS.
b. For heavy workload, VRQ is ~5 seconds(3%) faster than original BFS and the baseline code.
It's kind of busy these days I have tried some new idea based on VRQ solution, but the result isn't that good, so I spend most of time in debugging and haven't update the blog, :)
Based on the test result of VRQ, I believe it is a good direction to go. Especially it shows little overhead under heavy load than original BFS. And recently, I am thinking about improving BFS/VRQ by reducing grq usage, but the debug code prove an existed bug in current VRQ code, now I have to rework the whole VRQ code allover again to debug it. So in the coming 0.3 release of VRQ, I plan to fix the known issues in vrq branch and clean-up the code, which will help to stabilize VRQ code and also helpful for new idea which based on VRQ solution. At the meantime, I have update the -gc branch and put detail comments into bfs commits(if needed), hopefully CK will pick-up some of them in next BFS release.
Stay tuned. :D
#1 Jobs/Cores ratio 50%
3.17_0458_50
5m20.933s
5m20.896s
5m20.882s
5m20.867s
5m20.849s
5m20.743s
5m20.663s
5m20.621s
5m20.340s
5m20.251s
3.17_Baseline_50
5m18.065s
5m18.019s
5m17.956s
5m17.953s
5m17.941s
5m17.896s
5m17.891s
5m17.839s
5m17.826s
5m17.794s
3.17_VRQ_50
5m17.897s
5m17.856s
5m17.852s
5m17.788s
5m17.759s
5m17.748s
5m17.721s
5m17.673s
5m17.642s
5m17.634s
#2 Jobs/Cores ratio 150%
3.17_0458_150
2m56.468s
2m56.379s
2m56.377s
2m56.367s
2m56.364s
2m56.335s
2m56.322s
2m56.321s
2m56.261s
2m56.184s
3.17_Baseline_150
2m56.472s
2m56.407s
2m56.341s
2m56.312s
2m56.303s
2m56.293s
2m56.289s
2m56.288s
2m56.259s
2m56.202s
3.17_VRQ_150
2m51.553s
2m51.516s
2m51.473s
2m51.472s
2m51.448s
2m51.379s
2m51.358s
2m51.339s
2m51.338s
2m51.292s
#3 Jobs/Cores ratio 100%
3.17_0458_100
2m52.031s
2m50.718s
2m50.663s
2m50.650s
2m50.624s
2m50.603s
2m50.559s
2m50.541s
2m50.485s
2m50.468s
3.17_Baseline_100
2m50.945s
2m50.719s
2m50.683s
2m50.639s
2m50.602s
2m50.597s
2m50.583s
2m50.530s
2m50.521s
2m50.486s
3.17_VRQ_100
2m49.896s
2m49.675s
2m49.670s
2m49.658s
2m49.604s
2m49.538s
2m49.501s
2m49.480s
2m49.463s
2m49.457s
#4 Jobs/Cores ratio 100%
3.17_0458_100_50
2m52.026s
2m51.384s
2m51.151s
2m51.039s
2m50.983s
2m50.901s
2m50.863s
2m50.852s
2m50.810s
2m50.699s
3.17_Baseline_100_50
2m52.161s
2m51.007s
2m50.934s
2m50.864s
2m50.861s
2m50.857s
2m50.842s
2m50.817s
2m50.789s
2m50.733s
3.17_VRQ_100_50
2m49.801s
2m49.783s
2m49.740s
2m49.730s
2m49.716s
2m49.703s
2m49.700s
2m49.666s
2m49.662s
2m49.606s
As 0458 BFS released, I kicked off another new branch of testing to get the result of BFS 0458, Baseline of VRQ patch and the VRQ patch.
The result highlighted are
a. For low workload, Baseline&VRQ is ~3 seconds(1%) faster than original BFS.
b. For heavy workload, VRQ is ~5 seconds(3%) faster than original BFS and the baseline code.
It's kind of busy these days I have tried some new idea based on VRQ solution, but the result isn't that good, so I spend most of time in debugging and haven't update the blog, :)
Based on the test result of VRQ, I believe it is a good direction to go. Especially it shows little overhead under heavy load than original BFS. And recently, I am thinking about improving BFS/VRQ by reducing grq usage, but the debug code prove an existed bug in current VRQ code, now I have to rework the whole VRQ code allover again to debug it. So in the coming 0.3 release of VRQ, I plan to fix the known issues in vrq branch and clean-up the code, which will help to stabilize VRQ code and also helpful for new idea which based on VRQ solution. At the meantime, I have update the -gc branch and put detail comments into bfs commits(if needed), hopefully CK will pick-up some of them in next BFS release.
Stay tuned. :D
#1 Jobs/Cores ratio 50%
3.17_0458_50
5m20.933s
5m20.896s
5m20.882s
5m20.867s
5m20.849s
5m20.743s
5m20.663s
5m20.621s
5m20.340s
5m20.251s
3.17_Baseline_50
5m18.065s
5m18.019s
5m17.956s
5m17.953s
5m17.941s
5m17.896s
5m17.891s
5m17.839s
5m17.826s
5m17.794s
3.17_VRQ_50
5m17.897s
5m17.856s
5m17.852s
5m17.788s
5m17.759s
5m17.748s
5m17.721s
5m17.673s
5m17.642s
5m17.634s
#2 Jobs/Cores ratio 150%
3.17_0458_150
2m56.468s
2m56.379s
2m56.377s
2m56.367s
2m56.364s
2m56.335s
2m56.322s
2m56.321s
2m56.261s
2m56.184s
3.17_Baseline_150
2m56.472s
2m56.407s
2m56.341s
2m56.312s
2m56.303s
2m56.293s
2m56.289s
2m56.288s
2m56.259s
2m56.202s
3.17_VRQ_150
2m51.553s
2m51.516s
2m51.473s
2m51.472s
2m51.448s
2m51.379s
2m51.358s
2m51.339s
2m51.338s
2m51.292s
#3 Jobs/Cores ratio 100%
3.17_0458_100
2m52.031s
2m50.718s
2m50.663s
2m50.650s
2m50.624s
2m50.603s
2m50.559s
2m50.541s
2m50.485s
2m50.468s
3.17_Baseline_100
2m50.945s
2m50.719s
2m50.683s
2m50.639s
2m50.602s
2m50.597s
2m50.583s
2m50.530s
2m50.521s
2m50.486s
3.17_VRQ_100
2m49.896s
2m49.675s
2m49.670s
2m49.658s
2m49.604s
2m49.538s
2m49.501s
2m49.480s
2m49.463s
2m49.457s
#4 Jobs/Cores ratio 100%
3.17_0458_100_50
2m52.026s
2m51.384s
2m51.151s
2m51.039s
2m50.983s
2m50.901s
2m50.863s
2m50.852s
2m50.810s
2m50.699s
3.17_Baseline_100_50
2m52.161s
2m51.007s
2m50.934s
2m50.864s
2m50.861s
2m50.857s
2m50.842s
2m50.817s
2m50.789s
2m50.733s
3.17_VRQ_100_50
2m49.801s
2m49.783s
2m49.740s
2m49.730s
2m49.716s
2m49.703s
2m49.700s
2m49.666s
2m49.662s
2m49.606s
Tuesday, November 4, 2014
What's new for 3.17 gc patch set
In the last kernel release cycle, I played around some ideas with BFS, ideas like how to select a task to run, sticky task etc. The result is not as good as expected but all these trying let me knows BFS code better. So there will be no huge changes in BFS VRQ solution branch, I may need another round to think it all over again. Instead, there are some changes which can barely apply upon the original bfs grq lock solution are back-ported to -gc branch.
3e678e6 bfs: sync with mainline sched_setscheduler() logic.
-- There is new parameter called sched_attr is introduced during 3.14 in sched_setscheduler() and related functions. BFS is not fully sync with these changes yet.
f3a98f9 bfs: Refactory online_cpus() checking in try_preempt().
9c72372 bfs: Refactory needs_other_cpu().
--Two changes in try_preempt().
65f06e8 bfs: priodl to speed up try_preempt().
-- Introduced task priodl, which first 8 bits is the task prio and the last 56 bits are the task deadline(higher 56 bits) to speed up try_preempt() calculation.
1fa3bb5 bfs: Full cpumask based and LLC sensitive cpu selection.
-- Rewrite the cpu selection logic for tasks, by using full cpumask based calculation. The benefit of cpumask based calculation is that the cost is not scaled with cpu numbers when it is among a certain range(64 cpus for 64bits system and 32 cpus for 32bit system). The cost to transfer tasks among cpu which shares same LLC(Last Level Cache) should be consider free.
The best_mask_cpu() cpu selection logic now follows below orders:
* Non scaled same cpu as task originally runs on
* Non scaled SMT of the cup
* Non scaled cores/threads shares last level cache
* Scaled same cpu as task originally runs on
* Scaled same cpu as task originally runs on
* Scaled SMT of the cup
* Scaled cores/threads shares last level cache
* Non scaled cores within the same physical cpu
* Non scaled cpus/Cores within the local NODE
* Scaled cores within the same physical cpu
* Scaled cpus/Cores within the local NODE
* All cpus avariable
To implement full cpumask calculation, non_scaled_cpumask is introduced in grq structure. The plug-in code in cpufreq and intel_pstate drivers also be modified to avoid multi-trigger when scaling down from max cpu freq(intel_pstate driver just pass compile test, I have no hardware which runs on intel_pstate driver)
Here are the test result of the 3.17-gc, comparing to vrq-02-baseline-test-result, in low workload, these patches give another 3~4 seconds improvement.
#1 3.17.2-gc 50% tasks/cores ratio
5m18.630s
5m18.509s
5m18.504s
5m18.494s
5m18.489s
5m18.487s
5m18.481s
5m18.461s
5m18.383s
5m18.339s
#2 3.17.2-gc 150% tasks/cores ratio
2m56.475s
2m56.447s
2m56.418s
2m56.410s
2m56.402s
2m56.367s
2m56.349s
2m56.197s
2m56.173s
2m56.128s
#3 3.17.2-gc 100% tasks/cores ratio
2m52.318s
2m50.698s
2m50.654s
2m50.596s
2m50.578s
2m50.543s
2m50.534s
2m50.508s
2m50.437s
2m50.430s
#4 3.17.2-gc 100%+50% tasks/cores ratio
2m52.115s
2m51.246s
2m50.892s
2m50.880s
2m50.847s
2m50.819s
2m50.805s
2m50.804s
2m50.789s
2m50.645s
If you want to try these new patches, please check my linux-3.17.y-gc git branch.
3e678e6 bfs: sync with mainline sched_setscheduler() logic.
-- There is new parameter called sched_attr is introduced during 3.14 in sched_setscheduler() and related functions. BFS is not fully sync with these changes yet.
f3a98f9 bfs: Refactory online_cpus() checking in try_preempt().
9c72372 bfs: Refactory needs_other_cpu().
--Two changes in try_preempt().
65f06e8 bfs: priodl to speed up try_preempt().
-- Introduced task priodl, which first 8 bits is the task prio and the last 56 bits are the task deadline(higher 56 bits) to speed up try_preempt() calculation.
1fa3bb5 bfs: Full cpumask based and LLC sensitive cpu selection.
-- Rewrite the cpu selection logic for tasks, by using full cpumask based calculation. The benefit of cpumask based calculation is that the cost is not scaled with cpu numbers when it is among a certain range(64 cpus for 64bits system and 32 cpus for 32bit system). The cost to transfer tasks among cpu which shares same LLC(Last Level Cache) should be consider free.
The best_mask_cpu() cpu selection logic now follows below orders:
* Non scaled same cpu as task originally runs on
* Non scaled SMT of the cup
* Non scaled cores/threads shares last level cache
* Scaled same cpu as task originally runs on
* Scaled same cpu as task originally runs on
* Scaled SMT of the cup
* Scaled cores/threads shares last level cache
* Non scaled cores within the same physical cpu
* Non scaled cpus/Cores within the local NODE
* Scaled cores within the same physical cpu
* Scaled cpus/Cores within the local NODE
* All cpus avariable
To implement full cpumask calculation, non_scaled_cpumask is introduced in grq structure. The plug-in code in cpufreq and intel_pstate drivers also be modified to avoid multi-trigger when scaling down from max cpu freq(intel_pstate driver just pass compile test, I have no hardware which runs on intel_pstate driver)
Here are the test result of the 3.17-gc, comparing to vrq-02-baseline-test-result, in low workload, these patches give another 3~4 seconds improvement.
#1 3.17.2-gc 50% tasks/cores ratio
5m18.630s
5m18.509s
5m18.504s
5m18.494s
5m18.489s
5m18.487s
5m18.481s
5m18.461s
5m18.383s
5m18.339s
#2 3.17.2-gc 150% tasks/cores ratio
2m56.475s
2m56.447s
2m56.418s
2m56.410s
2m56.402s
2m56.367s
2m56.349s
2m56.197s
2m56.173s
2m56.128s
#3 3.17.2-gc 100% tasks/cores ratio
2m52.318s
2m50.698s
2m50.654s
2m50.596s
2m50.578s
2m50.543s
2m50.534s
2m50.508s
2m50.437s
2m50.430s
#4 3.17.2-gc 100%+50% tasks/cores ratio
2m52.115s
2m51.246s
2m50.892s
2m50.880s
2m50.847s
2m50.819s
2m50.805s
2m50.804s
2m50.789s
2m50.645s
If you want to try these new patches, please check my linux-3.17.y-gc git branch.
Monday, September 22, 2014
VRQ 0.2 release
As 3.17 will be released soon, earlier than it's expected, VRQ development is cut off and tagged for 0.2 release.
There are some bug fixes and others are improvement. Some is not related to VRQ locking, and I will see if it can be back-port to original BFS as baseline improvement in the next release. The detail changes are:
3ef882c bfs: Rework swap_sticky().
-- Yet another activity will be continued in next release.
e8754f9 bfs: rework resched_xxxx_idle(), basic version.
-- I will write another post to describe it in detail, but in brief, it rewrite the resched_xxxx_idle() using cpumask method.
2ae0fb6 bfs: refactory schedule() for rq&grq lock ctx switching.
cea6ce8 bfs: vrq: rq&grq locking ctx switch v3.
-- It's a bad idea to separate a context_switch process into two grq locking sessions, so I turn to this solution which hold rq and grq locking during context_switch.
319cd02 bfs: vrq, refactory wake_up_new_task.
79f5644 bfs: Fix need_other_cpu logic in schedule().
4d511fb bfs: RQ niffy solution.
-- Already described in previous post.
7c519a7 bfs: inlined routines update.
The test result are
50% Ratio:
5m19.531s
5m19.519s
5m19.509s
5m19.508s
5m19.430s
5m19.376s
5m19.363s
5m19.359s
5m19.333s
5m19.299s
150% Ratio:
2m54.394s
2m53.632s
2m51.960s
2m51.929s
2m51.925s
2m51.801s
2m51.790s
2m51.747s
2m51.641s
2m51.592s
100% Ratio:
2m51.001s
2m50.150s
2m49.881s
2m49.860s
2m49.812s
2m49.770s
2m49.764s
2m49.733s
2m49.699s
2m49.660s
100%+50% Ratio:
2m49.987s
2m49.980s
2m49.916s
2m49.865s
2m49.835s
2m49.828s
2m49.802s
2m49.784s
2m49.744s
2m49.733s
Comparing to vrq-02-baseline-test-result, under low or heavy workload, VRQ 0.2 shows a visible better throughput than the baseline. And under the optimize workload, VRQ 0.2 shows a slight better than baseline.
If you want have a try with VRQ 0.2, the code is located at v3.16.2-vrq.
There are some bug fixes and others are improvement. Some is not related to VRQ locking, and I will see if it can be back-port to original BFS as baseline improvement in the next release. The detail changes are:
3ef882c bfs: Rework swap_sticky().
-- Yet another activity will be continued in next release.
e8754f9 bfs: rework resched_xxxx_idle(), basic version.
-- I will write another post to describe it in detail, but in brief, it rewrite the resched_xxxx_idle() using cpumask method.
2ae0fb6 bfs: refactory schedule() for rq&grq lock ctx switching.
cea6ce8 bfs: vrq: rq&grq locking ctx switch v3.
-- It's a bad idea to separate a context_switch process into two grq locking sessions, so I turn to this solution which hold rq and grq locking during context_switch.
319cd02 bfs: vrq, refactory wake_up_new_task.
79f5644 bfs: Fix need_other_cpu logic in schedule().
4d511fb bfs: RQ niffy solution.
-- Already described in previous post.
7c519a7 bfs: inlined routines update.
The test result are
50% Ratio:
5m19.531s
5m19.519s
5m19.509s
5m19.508s
5m19.430s
5m19.376s
5m19.363s
5m19.359s
5m19.333s
5m19.299s
150% Ratio:
2m54.394s
2m53.632s
2m51.960s
2m51.929s
2m51.925s
2m51.801s
2m51.790s
2m51.747s
2m51.641s
2m51.592s
100% Ratio:
2m51.001s
2m50.150s
2m49.881s
2m49.860s
2m49.812s
2m49.770s
2m49.764s
2m49.733s
2m49.699s
2m49.660s
100%+50% Ratio:
2m49.987s
2m49.980s
2m49.916s
2m49.865s
2m49.835s
2m49.828s
2m49.802s
2m49.784s
2m49.744s
2m49.733s
Comparing to vrq-02-baseline-test-result, under low or heavy workload, VRQ 0.2 shows a visible better throughput than the baseline. And under the optimize workload, VRQ 0.2 shows a slight better than baseline.
If you want have a try with VRQ 0.2, the code is located at v3.16.2-vrq.
Wednesday, September 17, 2014
VRQ 0.2: RQ niffy solution
One of the changes in VRQ 0.2 is RQ niffy solution, which is a replacement solution of grq niffies by put niffy into each RQ. For the original design of grq niffies, please read CK's post http://ck-hack.blogspot.com/2010/10/of-jiffies-gjiffies-miffies-niffies-and.html
Functions which need grq.niffies are time_slice_expired() and task_prio().
There are update_clocks() called before every time_slice_expired() with grq lock, that means there is no impact if RQ niffy solution is used instead of grq.niffies solution in update_clocks() and time_slice_expired().
In task_prio(), grq.niffies can be replaced by niffy in current RQ, it may not be the latest niffy among all the RQs, but it is acceptable.
By using RQ niffy solution, grq lock for niffy update/read is not required. It is designed to reduce grq lock hot spots.
For the code change, please check https://bitbucket.org/alfredchen/linux-gc/commits/f6ec6f5303cb88e7462f4321b7a29d6c8ab83e89?at=linux-3.16.y-vrq
Functions which need grq.niffies are time_slice_expired() and task_prio().
There are update_clocks() called before every time_slice_expired() with grq lock, that means there is no impact if RQ niffy solution is used instead of grq.niffies solution in update_clocks() and time_slice_expired().
In task_prio(), grq.niffies can be replaced by niffy in current RQ, it may not be the latest niffy among all the RQs, but it is acceptable.
By using RQ niffy solution, grq lock for niffy update/read is not required. It is designed to reduce grq lock hot spots.
For the code change, please check https://bitbucket.org/alfredchen/linux-gc/commits/f6ec6f5303cb88e7462f4321b7a29d6c8ab83e89?at=linux-3.16.y-vrq
Subscribe to:
Posts (Atom)