Monday, 1 August 2016

assembly - Unexpectedly poor and weirdly bimodal performance for store loop on Intel Skylake

I'm seeing unexpectedly poor performance for a simple store loop which has two stores: one with a forward stride of 16 byte and one that's always to the same location1, like this:




volatile uint32_t value;

void weirdo_cpp(size_t iters, uint32_t* output) {

uint32_t x = value;
uint32_t *rdx = output;
volatile uint32_t *rsi = output;
do {
*rdx = x;

*rsi = x;

rdx += 4; // 16 byte stride
} while (--iters > 0);
}


In assembly this loop probably3 looks like:



weirdo_cpp:


...

align 16
.top:
mov [rdx], eax ; stride 16
mov [rsi], eax ; never changes

add rdx, 16


dec rdi
jne .top

ret


When the memory region accessed is in L2 I would expect this to run at less than 3 cycles per iteration. The second store just keeps hitting the same location and should add about a cycle. The first store implies bringing in a line from L2 and hence also evicting a line once every 4 iterations. I'm not sure how you evaluate the L2 cost, but even if you conservatively estimate that the L1 can only do one of the following every cycle: (a) commit a store or (b) receive a line from L2 or (c) evict a line to L2, you'd get something like 1 + 0.25 + 0.25 = 1.5 cycles for the stride-16 store stream.



Indeed, you comment out one store you get ~1.25 cycles per iteration for the first store only, and ~1.01 cycles per iteration for the second store, so 2.5 cycles per iteration seems like a conservative estimate.




The actual performance is very odd, however. Here's a typical run of the test harness:



Estimated CPU speed:  2.60 GHz
output size : 64 KiB
output alignment: 32
3.90 cycles/iter, 1.50 ns/iter, cpu before: 0, cpu after: 0
3.90 cycles/iter, 1.50 ns/iter, cpu before: 0, cpu after: 0
3.90 cycles/iter, 1.50 ns/iter, cpu before: 0, cpu after: 0
3.89 cycles/iter, 1.49 ns/iter, cpu before: 0, cpu after: 0
3.90 cycles/iter, 1.50 ns/iter, cpu before: 0, cpu after: 0

4.73 cycles/iter, 1.81 ns/iter, cpu before: 0, cpu after: 0
7.33 cycles/iter, 2.81 ns/iter, cpu before: 0, cpu after: 0
7.33 cycles/iter, 2.81 ns/iter, cpu before: 0, cpu after: 0
7.34 cycles/iter, 2.81 ns/iter, cpu before: 0, cpu after: 0
7.26 cycles/iter, 2.80 ns/iter, cpu before: 0, cpu after: 0
7.28 cycles/iter, 2.80 ns/iter, cpu before: 0, cpu after: 0
7.31 cycles/iter, 2.81 ns/iter, cpu before: 0, cpu after: 0
7.29 cycles/iter, 2.81 ns/iter, cpu before: 0, cpu after: 0
7.28 cycles/iter, 2.80 ns/iter, cpu before: 0, cpu after: 0
7.29 cycles/iter, 2.80 ns/iter, cpu before: 0, cpu after: 0

7.27 cycles/iter, 2.80 ns/iter, cpu before: 0, cpu after: 0
7.30 cycles/iter, 2.81 ns/iter, cpu before: 0, cpu after: 0
7.30 cycles/iter, 2.81 ns/iter, cpu before: 0, cpu after: 0
7.28 cycles/iter, 2.80 ns/iter, cpu before: 0, cpu after: 0
7.28 cycles/iter, 2.80 ns/iter, cpu before: 0, cpu after: 0


Two things are weird here.



First are the bimodal timings: there is a fast mode and a slow mode. We start out in slow mode taking about 7.3 cycles per iteration, and at some point transition to about 3.9 cycles per iteration. This behavior is consistent and reproducible and the two timings are always quite consistent clustered around the two values. The transition shows up in both directions from slow mode to fast mode and the other way around (and sometimes multiple transitions in one run).




The other weird thing is the really bad performance. Even in fast mode, at about 3.9 cycles the performance is much worse than the 1.0 + 1.3 = 2.3 cycles worst cast you'd expect from adding together the each of the cases with a single store (and assuming that absolutely zero worked can be overlapped when both stores are in the loop). In slow mode, performance is terrible compared to what you'd expect based on first principles: it is taking 7.3 cycles to do 2 stores, and if you put it in L2 store bandwidth terms, that's roughly 29 cycles per L2 store (since we only store one full cache line every 4 iterations).



Skylake is recorded as having a 64B/cycle throughput between L1 and L2, which is way higher than the observed throughput here (about 2 bytes/cycle in slow mode).



What explains the poor throughput and bimodal performance and can I avoid it?



I'm also curious if this reproduces on other architectures and even on other Skylake boxes. Feel free to include local results in the comments.



You can find the test code and harness on github. There is a Makefile for Linux or Unix-like platforms, but it should be relatively easy to build on Windows too. If you want to run the asm variant you'll need nasm or yasm for the assembly4 - if you don't have that you can just try the C++ version.




Eliminated Possibilities



Here are some possibilities that I considered and largely eliminated. Many of the possibilities are eliminated by the simple fact that you see the performance transition randomly in the middle of the benchmarking loop, when many things simply haven't changed (e.g., if it was related to the output array alignment, it couldn't change in the middle of a run since the same buffer is used the entire time). I'll refer to this as the default elimination below (even for things that are default elimination there is often another argument to be made).




  • Alignment factors: the output array is 16 byte aligned, and I've tried up to 2MB alignment without change. Also eliminated by the default elimination.

  • Contention with other processes on the machine: the effect is observed more or less identically on an idle machine and even on a heavily loaded one (e.g., using stress -vm 4). The benchmark itself should be completely core-local anyways since it fits in L2, and perf confirms there are very few L2 misses per iteration (about 1 miss every 300-400 iterations, probably related to the printf code).

  • TurboBoost: TurboBoost is completely disabled, confirmed by three different MHz readings.

  • Power-saving stuff: The performance governor is intel_pstate in performance mode. No frequency variations are observed during the test (CPU stays essentially locked at 2.59 GHz).


  • TLB effects: The effect is present even when the output buffer is located in a 2 MB huge page. In any case, the 64 4k TLB entries more than cover the 128K output buffer. perf doesn't report any particularly weird TLB behavior.

  • 4k aliasing: older, more complex versions of this benchmark did show some 4k aliasing but this has been eliminated since there are no loads in the benchmark (it's loads that might incorrectly alias earlier stores). Also eliminated by the default elimination.

  • L2 associativity conflicts: eliminated by the default elimination and by the fact that this doesn't go away even with 2MB pages, where we can be sure the output buffer is laid out linearly in physical memory.

  • Hyperthreading effects: HT is disabled.

  • Prefetching: Only two of the prefetchers could be involved here (the "DCU", aka L1<->L2 prefetchers), since all the data lives in L1 or L2, but the performance is the same with all prefetchers enabled or all disabled.

  • Interrupts: no correlation between interrupt count and slow mode. There is a limited number of total interrupts, mostly clock ticks.



toplev.py




I used toplev.py which implements Intel's Top Down analysis method, and to no surprise it identifies the benchmark as store bound:



BE             Backend_Bound:                                                      82.11 % Slots      [  4.83%]
BE/Mem Backend_Bound.Memory_Bound: 59.64 % Slots [ 4.83%]
BE/Core Backend_Bound.Core_Bound: 22.47 % Slots [ 4.83%]
BE/Mem Backend_Bound.Memory_Bound.L1_Bound: 0.03 % Stalls [ 4.92%]
This metric estimates how often the CPU was stalled without
loads missing the L1 data cache...
Sampling events: mem_load_retired.l1_hit:pp mem_load_retired.fb_hit:pp
BE/Mem Backend_Bound.Memory_Bound.Store_Bound: 74.91 % Stalls [ 4.96%] <==

This metric estimates how often CPU was stalled due to
store memory accesses...
Sampling events: mem_inst_retired.all_stores:pp
BE/Core Backend_Bound.Core_Bound.Ports_Utilization: 28.20 % Clocks [ 4.93%]
BE/Core Backend_Bound.Core_Bound.Ports_Utilization.1_Port_Utilized: 26.28 % CoreClocks [ 4.83%]
This metric represents Core cycles fraction where the CPU
executed total of 1 uop per cycle on all execution ports...
MUX: 4.65 %
PerfMon Event Multiplexing accuracy indicator



This doesn't really shed much light: we already knew it must be the stores messing things up, but why? Intel's description of the condition doesn't say much.



Here's a reasonable summary of some of the issues involved in L1-L2 interaction.






Update Feb 2019: I cannot no longer reproduce the "bimodal" part of the performance: for me, on the same i7-6700HQ box, the performance is now always very slow in the same cases the slow and very slow bimodal performance applies, i.e., with results around 16-20 cycles per line, like this:



Everything slow now




This change seems to have been introduced in the August 2018 Skylake microcode update, revision 0xC6. The prior microcode, 0xC2 shows the original behavior described in the question.






1 This is a greatly simplified MCVE of my original loop, which was at least 3 times the size and which did lots of additional work, but exhibited exactly the same performance as this simple version, bottlenecked on the same mysterious issue.



3 In particular, it looks exactly like this if you write the assembly by hand, or if you compile it with gcc -O1 (version 5.4.1), and probably most reasonable compilers (volatile is used to avoid sinking the mostly-dead second store outside the loop).



4 No doubt you could convert this to MASM syntax with a few minor edits since the assembly is so trivial. Pull requests accepted.

php - Not working my .htaccess and config files

I changed to the config file and .htaccess files. following files,



Config.php




$config['base_url'] = BASE_URL;

$config['index_page'] = '';
$config['uri_protocol'] = 'REQUEST_URI';



.htaccess




RewriteEngine On




RewriteBase /52322/



RewriteCond %{REQUEST_FILENAME} !-f



RewriteCond %{REQUEST_FILENAME} !-d



RewriteRule ^(.*)$ index.php?/$1 [L,QSA]




Error occur The requested URL /52322/login/ was not found on this server.




Please help me.



herewith I have attached for a screenshot.
enter image description here

objective c - Location Services not working in iOS 8



My app that worked fine on iOS 7 doesn't work with the iOS 8 SDK.



CLLocationManager doesn't return a location, and I don't see my app under Settings -> Location Services either. I did a Google search on the issue, but nothing came up. What could be wrong?



Answer



I ended up solving my own problem.



Apparently in iOS 8 SDK, requestAlwaysAuthorization (for background location) or requestWhenInUseAuthorization (location only when foreground) call on CLLocationManager is needed before starting location updates.



There also needs to be NSLocationAlwaysUsageDescription or NSLocationWhenInUseUsageDescription key in Info.plist with a message to be displayed in the prompt. Adding these solved my problem.



enter image description here



Hope it helps someone else.




EDIT: For more extensive information, have a look at: Core-Location-Manager-Changes-in-ios-8


c# - Having a collection in class




There are several options when one class must have a container (collection) of some sort of objects and I was wondering what implementation I shall prefer.




Here follow the options I found:



public class AClass : IEnumerable{
private List values = new List()

IEnumerator IEnumerable.GetEnumerator()
{
return GetEnumerator();
}


public IEnumerator GetEnumerator(){
return values.GetEnumerator();
}
}


Pros: AClass is not dependent on a concrete implementation of a collection (in this case List).



Cons: AClass doesn't have interface for Adding and removing elements




public class AClass : ICollection{
private List values = new List()

IEnumerator IEnumerable.GetEnumerator()
{
return GetEnumerator();
}

public IEnumerator GetEnumerator(){
return values.GetEnumerator();

}

//other ICollectionMembers
}


Pros: Same as IEnumerable plus it have interface for adding and removing elements



Cons: The ICollection interface define other methods that one rarely uses and it get's boring to implement those just for the sake of the interface. Also IEnumerable LINQ extensions takes care of some of those.




public class AClass : List{

}


Pros: No need of implementing any method. Caller may call any method implemented by List



Cons: AClass is dependent on collection List and if it changes some of the caller code may need to be changed. Also AClass can't inherit any other class.



The question is: Which one shall I prefer to state that my class contains a collection supporting both Add and Remove operations? Or other suggestions...



Answer



My suggestion is just define a generic List inside of your class and write additional Add and Remove methods like this and implement IEnumerable:



public class MyClass : IEnumerable
{
private List myList;

public MyClass()
{
myList = new List();

}

public void Add(string item)
{
if (item != null) myList.Add(item);
}

public void Remove(string item)
{
if (myList.IndexOf(item) > 0) myList.Remove(item);

}

public IEnumerable MyList { get { return myList; } }

public IEnumerator GetEnumerator()
{
return myList.GetEnumerator();
}
}



This is the best way if you don't want to implement your own collection.You don't need to implement an interface to Add and Remove methods.The additional methods like this fits your needs I guess.


php - Looping through a table with Simple HTML DOM



I'm using Simple HTML DOM to extract data from a HTML document, and I have a couple of issues that I need some help with.




  1. On the line that begins with if ($td->find('a')) I want to extract the href and the content of the anchor node separately, and place them in separate variables. The code however doesn't work (see output from echoes in the code below).



    What is the best way to do this? Note that my purpose is to create a XML document out of the information later on, so I need the information in the correct order.


  2. The links leads to pages containing detailed information about the different cars (e.g. "Max speed", "Price" etc) that I also want to extract and put into separate variables. How can I get hold of data on these pages?




    include 'simple_html_dom.php';

    $html = new simple_html_dom();
    $html = file_get_html('http://www.example.com/foo.html');

    $items = array();

    foreach ($html->find('table') as $table) {
    foreach ($table->find('tr') as $tr) {


    foreach ($tr->find('td') as $td) {

    if ($td->find('a')) {
    $link = $td->find('a.href');
    echo $link; // empty

    $text = $td->find('a.text');
    echo $text; // Array
    }

    else {
    echo 'Name: ' . $td;
    }
    }
    }
    }



The HTML document looks like this:















... and so on...


Answer



Use $td->find('a', 0)->href and $td->find('a', 0)->innertext to access element attributes in the first case, and contents in the second. Also, if there might be multiple anchor to be found, use 0 as a safe guard to always get the first one.


Hot Linked Questions

How does C handle EOF? [duplicate]




#include

int main()
{
FILE* f=fopen("book2.txt","r");
char a[200];
while(!feof(f))
{
fscanf(f,"%s",a);

printf("%s ",a);
printf("%d\n",ftell(f));...





PHP - Floating Number Precision





$a = '35';
$b = '-34.99';
echo ($a + $b);


Results in 0.009999999999998



What is up with that? I wondered why my program kept reporting odd results.



Why doesn't PHP return the expected 0.01?



Answer



Because floating point arithmetic != real number arithmetic. An illustration of the difference due to imprecision is, for some floats a and b, (a+b)-b != a. This applies to any language using floats.



Since floating point are binary numbers with finite precision, there's a finite amount of representable numbers, which leads accuracy problems and surprises like this. Here's another interesting read: What Every Computer Scientist Should Know About Floating-Point Arithmetic.






Back to your problem, basically there is no way to accurately represent 34.99 or 0.01 in binary (just like in decimal, 1/3 = 0.3333...), so approximations are used instead. To get around the problem, you can:





  1. Use round($result, 2) on the result to round it to 2 decimal places.


  2. Use integers. If that's currency, say US dollars, then store $35.00 as 3500 and $34.99 as 3499, then divide the result by 100.




It's a pity that PHP doesn't have a decimal datatype like other languages do.


c++ - Does curly brackets matter for empty constructor?

Those brackets declare an empty, inline constructor. In that case, with them, the constructor does exist, it merely does nothing more than t...


Car 1

Porsche

Car 2

Chrysler

Blog Archive