updated cuda todos. Please look at cuda-packet-batcher.c to have a look at the new todos

remotes/origin/master-1.1.x
Anoop Saldanha 16 years ago committed by Victor Julien
parent c734cd1bdd
commit 7dd2392dea

@ -4,19 +4,37 @@
* \author Anoop Saldanha <poonaatsoc@gmail.com> * \author Anoop Saldanha <poonaatsoc@gmail.com>
* *
* \todo * \todo
* - Make cuda paramters user configurable. * 1 Implement a gpu version of aho-corasick. That should get rid of a
* - Implement a gpu version of aho-corasick. That should get rid of a
* lot of post processing and pattern_chopping, and we don't have to * lot of post processing and pattern_chopping, and we don't have to
* deal with one or two byte patterns. * deal with one or two byte patterns. (currently in process)
* - Currently a lot of packets(~17k) are getting stuck on the detection * 2 Use texture/shared memory. This should be handled along with 1 and 6.
* 3 Currently a lot of packets(~17k) are getting stuck on the detection
* thread, which is a major bottleneck. Introduce bypass detection * thread, which is a major bottleneck. Introduce bypass detection
* threads for these 15k non buffered packets and check how the alerts * threads for these 15k non buffered packets and check how the alerts
* are affected by this(out of sequence handling by detection threads). * are affected by this(out of sequence handling by detection threads).
* - Use texture/shared memory. This should be handled along with AC. * 4 Test the use of mapped memory(if possible anywhere).
* - Test the use of host-alloced page locked memory. * 5 Check parallelising memcopies with kernel execution.
* - Test other optimizations like using the sgh held in the flow(if * 6 Test this feature - Rearrange the packet stream(either on cpu or gpu),
* present in the flow), instead of retrieving the sgh inside the batcher * where each block in the gpu can access the packet with non-coalesced
* thread. * reads.
* 2 packets p1 -> aabb ccdd
* p2 -> eeff gghh
*
* stream -> aabbeeffccddgghh.
*
* Modify the block size to 16 threads for CC < 2.0 devices and 32 for
* for >= 2.0.
*
* The rearrangement of packet stream can be done on the gpu, with no
* perf degradation, using coalesced reads. Padding packets need to
* be addressed though.
* (Need to give more thought to this task).
*
* -- Feel free to pick any task from the agenda, but please
* drop a mail to dev mailing list(or directly to the dev team). Better
* yet, open a feature request on our bug/feature tracker
* (https://redmine.openinfosecfoundation.org/issues). Will be a mess if
* 2 or more devs end up working on the same task or related tasks.
*/ */
/* compile in, only if we have a CUDA enabled on this machine */ /* compile in, only if we have a CUDA enabled on this machine */

Loading…
Cancel
Save