From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mga11.intel.com (mga11.intel.com [192.55.52.93]) by dpdk.org (Postfix) with ESMTP id C9CE2106A for ; Tue, 17 Jan 2017 07:25:06 +0100 (CET) Received: from orsmga003.jf.intel.com ([10.7.209.27]) by fmsmga102.fm.intel.com with ESMTP; 16 Jan 2017 22:25:05 -0800 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.33,243,1477983600"; d="scan'208";a="923358080" Received: from fmsmsx107.amr.corp.intel.com ([10.18.124.205]) by orsmga003.jf.intel.com with ESMTP; 16 Jan 2017 22:25:05 -0800 Received: from fmsmsx121.amr.corp.intel.com (10.18.125.36) by fmsmsx107.amr.corp.intel.com (10.18.124.205) with Microsoft SMTP Server (TLS) id 14.3.248.2; Mon, 16 Jan 2017 22:25:05 -0800 Received: from bgsmsx153.gar.corp.intel.com (10.224.23.4) by fmsmsx121.amr.corp.intel.com (10.18.125.36) with Microsoft SMTP Server (TLS) id 14.3.248.2; Mon, 16 Jan 2017 22:25:05 -0800 Received: from bgsmsx101.gar.corp.intel.com ([169.254.1.43]) by BGSMSX153.gar.corp.intel.com ([169.254.2.218]) with mapi id 14.03.0248.002; Tue, 17 Jan 2017 11:54:57 +0530 From: "Yang, Zhiyong" To: "thomas.monjalon@6wind.com" , "Richardson, Bruce" , "Ananyev, Konstantin" CC: "yuanhan.liu@linux.intel.com" , "De Lara Guarch, Pablo" , "dev@dpdk.org" Thread-Topic: [dpdk-dev] [PATCH v2 0/4] eal/common: introduce rte_memset and related test Thread-Index: AQHSal3HJXuT/EIHqUKnUx4wYUps0aE8Dzcw Date: Tue, 17 Jan 2017 06:24:55 +0000 Message-ID: References: <1480926387-63838-2-git-send-email-zhiyong.yang@intel.com> <1482833098-38096-1-git-send-email-zhiyong.yang@intel.com> In-Reply-To: Accept-Language: en-US Content-Language: en-US X-MS-Has-Attach: X-MS-TNEF-Correlator: x-titus-metadata-40: eyJDYXRlZ29yeUxhYmVscyI6IiIsIk1ldGFkYXRhIjp7Im5zIjoiaHR0cDpcL1wvd3d3LnRpdHVzLmNvbVwvbnNcL0ludGVsMyIsImlkIjoiOTlmNzFkMzYtODA1NS00MWJjLTlkNzktNTM3M2EwMDgwMzk5IiwicHJvcHMiOlt7Im4iOiJDVFBDbGFzc2lmaWNhdGlvbiIsInZhbHMiOlt7InZhbHVlIjoiQ1RQX0lDIn1dfV19LCJTdWJqZWN0TGFiZWxzIjpbXSwiVE1DVmVyc2lvbiI6IjE1LjkuNi42IiwiVHJ1c3RlZExhYmVsSGFzaCI6Im02OVArd0xEaEJGV0dUbjl1aXVwQWh0ZDhJY202dVo5c3ZpY1NjcmRBbmM9In0= x-ctpclassification: CTP_IC x-originating-ip: [10.223.10.10] Content-Type: text/plain; charset="us-ascii" Content-Transfer-Encoding: quoted-printable MIME-Version: 1.0 Subject: Re: [dpdk-dev] [PATCH v2 0/4] eal/common: introduce rte_memset and related test X-BeenThere: dev@dpdk.org X-Mailman-Version: 2.1.15 Precedence: list List-Id: DPDK patches and discussions List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , X-List-Received-Date: Tue, 17 Jan 2017 06:25:07 -0000 Hi, Thomas: Does this patchset have chance to be applied for 1702 release?=20 Thanks Zhiyong > -----Original Message----- > From: dev [mailto:dev-bounces@dpdk.org] On Behalf Of Yang, Zhiyong > Sent: Monday, January 9, 2017 5:49 PM > To: thomas.monjalon@6wind.com; Richardson, Bruce > ; Ananyev, Konstantin > > Cc: yuanhan.liu@linux.intel.com; De Lara Guarch, Pablo > ; dev@dpdk.org > Subject: Re: [dpdk-dev] [PATCH v2 0/4] eal/common: introduce rte_memset > and related test >=20 > Hi, Thomas, Bruce, Konstantin: >=20 > Any comments about the patchset? Do I need to modify anything? >=20 > Thanks > Zhiyong >=20 > > -----Original Message----- > > From: dev [mailto:dev-bounces@dpdk.org] On Behalf Of Zhiyong Yang > > Sent: Tuesday, December 27, 2016 6:05 PM > > To: dev@dpdk.org > > Cc: yuanhan.liu@linux.intel.com; thomas.monjalon@6wind.com; > > Richardson, Bruce ; Ananyev, Konstantin > > ; De Lara Guarch, Pablo > > > > Subject: [dpdk-dev] [PATCH v2 0/4] eal/common: introduce rte_memset > > and related test > > > > DPDK code has met performance drop badly in some case when calling > > glibc function memset. Reference to discussions about memset in > > http://dpdk.org/ml/archives/dev/2016-October/048628.html > > It is necessary to introduce more high efficient function to fix it. > > One important thing about rte_memset is that we can get clear control > > on what instruction flow is used. > > > > This patchset introduces rte_memset to bring more high efficient > > implementation, and will bring obvious perf improvement, especially > > for small N bytes in the most application scenarios. > > > > Patch 1 implements rte_memset in the file rte_memset.h on IA platform > > The file supports three types of instruction sets including sse & avx > > (128bits), > > avx2(256bits) and avx512(512bits). rte_memset makes use of > > vectorization and inline function to improve the perf on IA. In > > addition, cache line and memory alignment are fully taken into > consideration. > > > > Patch 2 implements functional autotest to validates the function > > whether to work in a right way. > > > > Patch 3 implements performance autotest separately in cache and memory. > > We can see the perf of rte_memset is obviously better than glibc > > memset especially for small N bytes. > > > > Patch 4 Using rte_memset instead of copy_virtio_net_hdr can bring > > 3%~4% performance improvements on IA platform from virtio/vhost > > non-mergeable loopback testing. > > > > Changes in V2: > > > > Patch 1: > > Rename rte_memset.h -> rte_memset_64.h and create a file > rte_memset.h > > for each arch. > > > > Patch 3: > > add the perf comparation data between rte_memset and memset on > > haswell. > > > > Patch 4: > > Modify release_17_02.rst description. > > > > Zhiyong Yang (4): > > eal/common: introduce rte_memset on IA platform > > app/test: add functional autotest for rte_memset > > app/test: add performance autotest for rte_memset > > lib/librte_vhost: improve vhost perf using rte_memset > > > > app/test/Makefile | 3 + > > app/test/test_memset.c | 158 +++++++++ > > app/test/test_memset_perf.c | 348 +++++++++++++= ++++++ > > doc/guides/rel_notes/release_17_02.rst | 7 + > > .../common/include/arch/arm/rte_memset.h | 36 ++ > > .../common/include/arch/ppc_64/rte_memset.h | 36 ++ > > .../common/include/arch/tile/rte_memset.h | 36 ++ > > .../common/include/arch/x86/rte_memset.h | 51 +++ > > .../common/include/arch/x86/rte_memset_64.h | 378 > > +++++++++++++++++++++ > > lib/librte_eal/common/include/generic/rte_memset.h | 52 +++ > > lib/librte_vhost/virtio_net.c | 18 +- > > 11 files changed, 1116 insertions(+), 7 deletions(-) create mode > > 100644 app/test/test_memset.c create mode 100644 > > app/test/test_memset_perf.c create mode 100644 > > lib/librte_eal/common/include/arch/arm/rte_memset.h > > create mode 100644 > > lib/librte_eal/common/include/arch/ppc_64/rte_memset.h > > create mode 100644 > > lib/librte_eal/common/include/arch/tile/rte_memset.h > > create mode 100644 > > lib/librte_eal/common/include/arch/x86/rte_memset.h > > create mode 100644 > > lib/librte_eal/common/include/arch/x86/rte_memset_64.h > > create mode 100644 > lib/librte_eal/common/include/generic/rte_memset.h > > > > -- > > 2.7.4