Dodo6/Topic_Modelling_using_LDA
1
1Article2 3Design and Analysis of an Approximate Adder with4Hybrid Error Reduction5Hyoju Seo 1, Yoon Seok Yang 2 and Yongtae Kim 1,*6School of Computer Science and Engineering, Kyungpook National University, Daegu 41566, Korea;7hyoju@knu.ac.kr82 Intel Labs, Intel Corporation, Santa Clara, CA 95054, USA; yoonseok.yang@intel.com9* Correspondence: yongtae@knu.ac.kr; Tel.: +82-53-750-757410111 12Received: 1 February 2020; Accepted: 10 March 2020; Published: 11 March 202013 14Abstract: This paper presents an energy-efficient approximate adder with a novel hybrid error15reduction scheme to significantly improve the computation accuracy at the cost of extremely low16additional power and area overheads. The proposed hybrid error reduction scheme utilizes only17two input bits and adjusts the approximate outputs to reduce the error distance, which leads to an18overall improvement in accuracy. The proposed design, when implemented in 65-nm CMOS19technology, has 3, 2, and 2 times greater energy, power, and area efficiencies, respectively, than20conventional accurate adders. In terms of the accuracy, the proposed hybrid error reduction scheme21allows that the error rate of the proposed adder decreases to 50% whereas those of the lower-part22OR adder and optimized lower-part OR constant adder reach 68% and 85%, respectively.23Furthermore, the proposed adder has up to 2.24, 2.24, and 1.16 times better performance with24respect to the mean error distance, normalized mean error distance (NMED), and mean relative25error distance, respectively, than the other approximate adder considered in this paper.26Importantly, because of an excellent design tradeoff among delay, power, energy, and accuracy, the27proposed adder is found to be the most competitive approximate adder when jointly analyzed in28terms of the hardware cost and computation accuracy. Specifically, our proposed adder achieves2951%, 49%, and 47% reductions of the power-, energy-, and error-delay-product-NMED products,30respectively, compared to the other considered approximate adders.31Keywords: approximate adder; approximate computing; hybrid error reduction; low power;32energy efficiency33 341. Introduction35Energy efficiency has become a critical requirement in the design of modern computing systems36and system-on-chips, and chip designers are being continually urged to develop energy-efficient37design techniques to meet this requirement. Approximate computing can offer remarkable energy38savings by trading-off accuracy [1]. This approach is based on the observation that not all applications39require 100% computation accuracy. Specifically, many digital signal processing (DSP) applications40are inherently error-resilient [2–4]. For example, humans may not recognize sporadic errors in digital41image processing, such as lossy discrete cosine transform, since they are usually negligible because42of human sensory limitations.43While approximate computing can be performed in all computing layers, ranging from software44to circuit level [5,6], in this paper, we focus on approximate circuits, particularly an approximate45adder. Certainly, an approximate adder is a fundamental arithmetic unit that is frequently used in46many error-tolerant applications to decrease the overall energy consumption, and it has drawn much47attention from researchers. As a result, numerous approximate adders have been proposed in the48literature [7–31], which will be reviewed in Section 2.49Electronics 2020, 9, 471; doi:10.3390/electronics903047150 51www.mdpi.com/journal/electronics52 53Electronics 2020, 9, 47154 552 of 1356 57In this paper, we propose a novel approximate adder with a hybrid error reduction technique58and systematically analyze and extensively compare our proposed adder with other approximate59adders in terms of the hardware performance and computation accuracy. The proposed scheme60improves the error rate (ER) of the lower-part OR adder (LOA) and error tolerant adder I (ETAI) from6168% to 50%. When the proposed adder is implemented in a 65-nm CMOS technology, its mean error62distance (MED) and mean relative error distance (MRED) performances are also enhanced by more63than 120% and more than 50%, respectively, compared to those of the ETAI at the cost of extremely64low area and power overheads (<10%). Impressively, our adder shows an excellent tradeoff between65the hardware cost and the computation accuracy and is clearly superior to the other considered66approximate adders. Specifically, the power- and energy-normalized mean error distance (NMED)67products of our approximate adder are up to 2.06 and 1.97 times, respectively, better than those of68the other considered approximate adders.69The remainder of this paper is organized as follows. In Section 2, we introduce some of the70related works that were recently published in the field of approximate adders. Then, we present the71proposed approximate adder by providing illustrative examples in Section 3. The hardware72architecture and error analysis of the adder are given in Section 3 as well. In Section 4, the results of73hardware implementation together with the systematic analysis and extensive comparison with the74other seven approximate adders in terms of the performance and accuracy are presented. In addition,75the joint analysis of the approximate adders is provided in Section 4. Finally, we conclude the work76in Section 5.772. Related Works78Lu proposed an approximate adder wherein the carry for each sum bit is predicted by the limited79number of its less significant bits to improve the overall speed [7]. Verma et al. reduced the area80overheads of Lu’s adder by sharing some components [8]. These adders are faster than the81conventional accurate adders because of the shorter carry propagation chain.82An equal-segment-based approximate adder splits an n-bit adder into several equally83partitioned smaller k-bit sub-adders that concurrently perform partial additions and partial carry84generations using a limited number of input bits. It should be noted that sub-adders can be85implemented in any type of conventional accurate adder, such as the ripple carry adder (RCA) and86carry lookahead adder (CLA). The adders proposed in [9] and [10] utilize the previous k input bits87only to predict the carry for sub-adders. Kim et al. [11,12] proposed an improved carry speculation88scheme for sub-adders that leverages 2k and more input bits to generate carries, which leads to a89much better computation accuracy than that of the adders in [9] and [10]. The accuracy-configurable90adder proposed by Kahng et al. [13] includes multiple 2k-bit sub-adders and makes k bits overlap to91generate approximate outputs.92In addition, approximate 1-bit full adders are utilized to add some least significant bits (LSBs)93in a multibit adder. For example, the LOA, as shown in Figure 1, employs OR gates to approximately94add several LSB inputs [14]. It divides an n-bit adder into a k-bit precise adder and an (n−k)-bit95approximate adder. An AND operation of two (n−k−1)th LSB inputs is performed to predict the carryin signal of the precise adder (i.e., Cin), and the precise adder takes the most significant k bit (MSB)96inputs and the carry to generate accurate outputs for the MSBs. The (n−k)-bit approximate adder is97realized via bit-by-bit OR operations. The adder proposed by Albicoocco et al. [15], termed the98LOAWA (i.e., LOA without AND operation) also uses the OR operation to realize the (n−k)-bit99approximate adder for the LSBs; however, the main difference with the LOA is the exclusion of the100carry prediction scheme to reduce hardware cost at the expense of accuracy. The optimized lowerpart constant OR adder (OLOCA) [16] is almost identical to the LOA. The approximate adder part101utilizes the OR function but a few LSB outputs are forced to “1” to reduce hardware cost. In other102words, only a few of the upper (n−k) LSB outputs are generated by the OR operation, and the103remaining lower bits of the LSB outputs are fixed to “1.” The hardware optimized error reduced104adder (HOERRA) proposed by Balasubramanian et al. [31] is another optimized version of the LOA,105which is suitable for both field programmable gate array and application specific integrated circuit-106 107Electronics 2020, 9, 471108 1093 of 13110 111based implementations. The approximate adder part of the HOERAA is similar to that of the OLOCA112in that the (n−k−2) LSB outputs are set to “1.” The remaining two bits of the part leverages OR113function, and a 2-to-1 multiplexer is used to produce (n−k−1)th LSB output by selecting either an OR114operation of two (n−k−1)th LSB inputs or an AND operation of two (n−k−2)th LSB inputs. The carryin signal (i.e., Cin) serves as the selection input of the multiplexer.115 116An-1:n-k Bn-1:n-k117 118An-k-1 An-k-1 An-k-2119Bn-k-1 Bn-k-1 Bn-k-2120 121A0122B0123 124k-bit125Precise Adder126Cin127Sn-1:n-k128 129Sn-k-1:0130 131Figure 1. Architecture of lower-part OR adder (LOA).132 133Similar to the LOA, the ETAI uses a k-bit precise adder and an (n−k)-bit approximate adder, the134latter of which is realized via its own modified XOR function [17] and is composed of a control block135and a carry-free addition block, as shown in Figure 2. It is important to note that the ETAI does not136include a carry prediction scheme for the precise adder and the carry is set to “0” (i.e., Cin = 0), which137leads to poor overall computation accuracy compared to that of the LOA. To improve the accuracy,138Kim developed the carry predicting ETA (CPETA) and the enhanced CPETA (ECPETA) in [18] and139[19], respectively, by adding low-cost carry prediction techniques to the original ETAI architecture.140Specifically, the CPETA uses the same AND-operation-based carry prediction scheme with the LOA,141and the ECPETA utilizes both the (n−k−1)th and (n−k−2)th LSB inputs with an additional OR gate to142predict the carry to improve the accuracy. The LOA, ETAI, and their variants show a good tradeoff143between the computation accuracy and the power and area costs thanks to the simplification of the144LSB additions.145 146An-1:n-k Bn-1:n-k147 148An-k-1:0149 150Bn-k-1:0151 152Control Block153k-bit154Precise Adder155 156Sn-1:n-k157 1580159 160Carry-Free161Addition Block162Sn-k-1:0163 164Figure 2. Architecture of error tolerant adder I (ETAI).165 166In addition to design of the approximate adder, the evaluation of its accuracy is also an167important task; some error metrics have been introduced for this purpose [32], such as the MED,168MRED, and NMED. These metrics are widely adopted along with power, energy, delay, and area to169evaluate and compare the performances of approximate adders [33].170 171Electronics 2020, 9, 471172 1734 of 13174 1753. Proposed Approximate Adder176In this section, we present our proposed approximate adder that enhances the computation177accuracy of the LOA by application of a novel hybrid error reduction scheme, and the resultant adder178is termed a hybrid error reduction LOA (HERLOA). We provide illustrative examples to effectively179introduce the proposed adder and use the following notations. Let An−1:0, Bn−1:0, and Sn−1:0 denote,180respectively, the two n-bit inputs and one n-bit output of the adder, and let S’n−1:0 denote the n-bit181intermediate approximate output prior to error reduction. Additionally, Ai, Bi, Si, and S’i denote the182(i)th LSBs of An−1:0, Bn−1:0, Sn−1:0, and S’n−1:0, respectively.1833.1. Operation of the Proposed Adder184Figure 3 shows the operation of the proposed adder with 16-bit inputs. An n-bit addition185requires two different kinds of additions. The precise part is obtained using a k-bit accurate adder186(e.g., RCA or CLA) for k MSBs, and the approximate part is obtained via OR and XOR operations of187the remaining (n−k) LSBs, where k < n. The example in Figure 3 splits a 16-bit input into an 8-bit188precise part and an 8-bit approximate part. It is worth mentioning that the bit-widths of two parts do189not need to be equal. The precise part performs an accurate addition using k MSB inputs, and the190carry-in signal (Cin) is predicted via an AND operation of two (n−k−1)th inputs (i.e., Cin = An−k−1 AND191Bn−k−1). Similar to the LOA, the approximate part leverages the OR function to add the lower order192input bits of two operands; however, the main difference with the LOA is that, in the proposed adder,193the bit operation of (n−k−1)th inputs is replaced with the XOR operation. This results in the formation194of a half adder for (n−k−1)th LSB inputs and effectively extends the length of accurate addition by one195bit, which consequently leads to an overall improvement in accuracy compared to that of the LOA.196In other words, the proposed n-bit adder always generates correct outputs for bit positions from (n−1)197to (n−k−1) by replacing the OR gate with the XOR gate at the (n−k−1)th LSB position. As a result, under198the input shown in Figure 3, the proposed adder generates the approximate part output of199“01011010”, whereas the LOA generates an output of “11011010.” Obviously, the output of the200proposed adder is closer to the correct summation of “01101010” than the output of the LOA, and the201error distance—which is defined as Sapproximate Saccurate , where Sapproximate and Saccurate are the202approximate and correct outputs, respectively—for the given input decreases from “1110000” (112)203to “10000” (16).204 205Precise Part206 207Approximate Part208 209MSB210 211LSB212 213An-1:0214 21510110111216 217Bn-1:0218 21900101001220 221Cin2221223 2241 10100102251 0011000226XOR227 228Sn-1:0229 23011100001231 2320 1011010233 234Figure 3. Operation of proposed adder.235 2363.2. Proposed Hybrid Error Reduction Scheme237The proposed approximate adder performs an error reduction to further decrease the output238errors when both (n−k−2)th input bits are “1” (i.e., An−k−2 = 1 and Bn−k−2 = 1). Otherwise, it does not239perform any error reduction, as shown in the example in Figure 3. In fact, our adder implements the240error reduction logic, but this does not affect the final outputs of the adder in this case. The proposed241hybrid error reduction is performed differently depending on the (n−k−1)th inputs. In other words,242the proposed adder checks the (n−k−1)th output bit to determine which of the two error reduction243schemes is applicable. If both the inputs are “1” or “0,” the reduction logic corrects the (n−k−1)th and244 245Electronics 2020, 9, 471246 2475 of 13248 249(n−k−2)th outputs to “1” and “0,” respectively (i.e., Sn−k−1:n−k−2 = “10”). Otherwise, it sets all the outputs250from Sn−k−3 to S0 as “1.”251The example inputs shown in Figure 4a yield the approximate part output of “01011010” after252the normal addition described in Section 3.1; we term this output the intermediate approximate253output. Then, the adder further checks the (n−k−1)th output bit because both the (n−k−2)th input bits254are “1”. An intermediate approximate output S’n−k−1 of “0” implies that the corresponding input bits255are identical because of the XOR operation performed at the (n−k−1) bit position, and the final adder256outputs at the (n−k−1) to (n−k−2) bit positions are “1” and “0”, respectively. In short, Sn−k−1:n−k−2 = “10,”257and this is, in fact, the correct summation of outputs at the corresponding positions under the given258inputs. This scheme leads to a 2n−k−2 reduction in the error distance.259Precise Part260 261Approximate Part262 263MSB264 265LSB266 267An-1:0268 26910110111270 271Bn-1:0272 27300101001274 275Cin2761277 2781279 2801010010281 2821283 2841011000285 286XOR287OR288 289S'n-1:0290 29111100001292 2930294 2951011010296 297Sn-1:0298 29911100001300 3011302 3030011010304 305(a)306Precise Part307 308Approximate Part309 310MSB311 312An-1:0313 31410110111315 316Bn-1:0317 31800101001319 320LSB321Cin3220323 3240 1 0100103251 1 011000326XOR327OR328 329S'n-1:0330 33111100000332 3331 1 011010334 335Sn-1:0336 33711100000338 3391 1 111111340(b)341 342Figure 4. Operations of proposed error reduction when (a) An−k−1 and Bn−k−1 are identical and (b) An−k−1343and Bn−k−1 are exclusive.344 345On the other hand, if the intermediate approximate output S’n−k−1 is “1”, as illustrated in Figure3464b, all the remaining lower order bits are forced to “1,” which results in the approximate output of347“11111111.” Under this given input condition, a carry is supposed to be generated in the (n−k−2)th348LSB and propagate to the precise part through the (n−k−1)th LSB. However, as observed in Figure 4b,349the carry does not actually propagate to the precise adder (i.e., Cin = 0) because the carry prediction is350performed using an AND operation with only the (n−k−1)th LSB inputs. This result means that the351correct summation will always be larger than the proposed addition in this case. Therefore, forcing352all the outputs of the approximate part to “1” brings the approximate output closer to the correct353summation. This reduction scheme enables up to a 2 n−k−2 − 1 decrease in the error distance.354 355Electronics 2020, 9, 471356 3576 of 13358 3593.3. Implementation of the Proposed Adder360Figure 5 shows the hardware implementation of the approximate part of the proposed361HERLOA. It should be noted that the precise part is the same as that in Figure 1, and Cin is fed to the362precise adder. In the approximate addition operation, the XOR and OR gates are used for the363(n−k−1)th LSB and the other lower order LSBs, respectively, to generate the intermediate approximate364outputs S’n−k−1:0. Furthermore, the output of the AND operation of the (n−k−2)th LSB determines365whether or not error reduction should be performed. The intermediate approximate outputs S’n−k−1:0366are fed to the INV, AND, OR, and NAND gates to compute the final approximate outputs Sn−k−1:0. It367is important to note that the intermediate approximate outputs will bypass the reduction logic when368neither of the (n−k−2)th inputs is “1.” The critical path delay of the approximate part, tapproximate, can be369expressed simply as follows:370 371tapproximate t XOR t INV t NAND t AND ,372 373(1)374 375where tINV, tXOR, tNAND, and tAND are the delays of an inverter, a two-input XOR gate, a NAND gate,376and an AND gate, respectively.377 378An-k-1 An-k-1 An-k-2379Bn-k-1 Bn-k-1 Bn-k-2380 381S'n-k-1382 383An-k-2 An-k-3384Bn-k-2 Bn-k-3385 386A0387B0388 389S'n-k-2390 391S'0392 393S'n-k-3394 395Cin396 397Sn-k-1398 399Sn-k-2 Sn-k-3400 401S0402 403Figure 5. Architecture of approximate part of proposed adder, i.e., hybrid error reduction LOA404(HERLOA).405 4063.4. Error Rate Analysis407An output error occurs when any bit positions from (n−k−3) to 0 of both the inputs, A and B, are408“1”, which generates a carry for the higher bit position. Moreover, an error occurs when both the409(n−k−2)th LSB inputs are “1” and the two (n−k−1)th LSB inputs are exclusive, as observed in Figure4104b. Therefore, the ER of the proposed HERLOA with the error reduction scheme under random input411patterns is given as follows:412 41373414ERHERLOA ( n, k ) 1 41584416 417nk 2418 419,420 421(2)422 423where n and k are the sizes of the entire adder and the precise adder, respectively.4244. Experimental Results425The proposed approximate adder with design parameters of n = 16 and k = 8 was designed in426Verilog HDL and was synthesized using 65-nm CMOS technology and a standard cell library to427evaluate the delay, area, power, power-delay product (PDP), and energy-delay product (EDP). The4288-bit RCA was employed as the precise adder. For comparison of our proposed adder with other429adders, two conventional accurate adders (RCA and CLA) as well as seven approximate adders430(LOA, ETAI, and their variants: LOAWA, OLOCA, HOERAA, CPETA, and ECPETA) were designed431 432Electronics 2020, 9, 471433 4347 of 13435 436and also synthesized using the same technology and library. For fair comparison, identical design437parameters, i.e., n = 16 and k = 8, and the RCA structure were used for all the approximate adders.438Furthermore, the bit-width of the constant part of the OLOCA and HOERAA were selected to be 6439[16,31].440In addition to implementing hardware, we constructed a software simulator to assess the441accuracy performance of the approximate adders in terms of the ER, MED, NMED, and MRED. These442error metrics are expressed as follows:443 444MED 445 446(3)447 448EDi4491 n450,451452n i 1 Si , accurate453 454(4)455 456MED 1 n EDi457,458 459D460n i 1 D461 462(5)463 464MRED 465 466NMED 467 4681 n469 EDi ,470n i 1471 472where n is the number of inputs, EDi is the error distance for the (i)th item of input data, Si,accurate is the473accurate output for the (i)th item of input data, and D is the maximum possible error value of the474approximate adder. These error metrics were estimated by using two samples each comprising 10475million (i.e., 107) uniformly distributed random input numbers.4764.1. Performance Analysis477Table 1 summarizes the performance of the proposed approximate adder and those of the other478eight adders. While the CLA is the fastest, the RCA has the longest delay because of the bit-by-bit479carry propagation. This delay of the RCA causes it to consume the largest amount of energy (i.e., the480highest PDP) even though its power dissipation is lower than that of the CLA, which consumes the481second largest amount of energy. The LOA and its variants (i.e., LOAWA, OLOCA, and HOERAA)482are more area-, power-, energy-, and EDP-efficient than the ETAI and its variants (i.e., CPETA and483ECPETA) because of the relatively simpler approximation scheme (i.e., OR operation) for the lower484half of the input bits in the case of the former category. Specifically, the OLOCA occupies the smallest485area, and the LOAWA is the fastest as well as the most power- and energy-efficient among all the486approximate adders. The simple AND-operation-based carry prediction in the LOA, OLOCA,487HOERAA, CPETA, and our proposed adder results in a longer delay than those of the ETAI and488LOAWA, which lack any carry prediction for the precise part. It should be noted that the critical path489delay of these adders exists in the precise adder part, including the carry prediction (i.e., the 8-bit490RCA with the AND gate). The delay of the HOERAA is insignificantly longer than that of the LOA,491OLOCA, CPETA, and the proposed adder even though all these adders have the identical carry492prediction scheme. It is because the carry prediction output (i.e., the AND gate output) of the493HOERAA is fed into not only the precise adder but also the multiplexer. This causes a higher fan-out494of the AND gate and therefore impacts on the delay. The more complicated carry prediction scheme495adopted in the ECPETA, which utilizes two MSB inputs from the approximate part, leads to the worst496area, delay, and power performances; as a result, this adder consumes the largest amount of energy497and has the highest EDP among the seven approximate adders. The proposed HERLOA is498comparable to the ETAI in terms of all the hardware performance metrics. It occupies 8% larger area499and consumes 9% more power than the LOA while having the same speed. A comparison of our500adder with the accurate adders reveals that our adder has up to 2.18, 1.96, 2.19, and 3.31 times greater501efficiency in terms of area, delay, power, and energy, respectively.502Our adder shows the best ER performance, whereas the OLOCA shows the worst ER503performance: it reaches over 99%, and the HOERAA has almost the same ER with the OLOCA. The504ERs of the LOA, LOAWA, and ETAI are the same and the ER of the CPETA is identical to that of the505 506Electronics 2020, 9, 471507 5088 of 13509 510ECPETA. The ETAI variants show better accuracy performance than the LOA and its variants in511terms of all the error metrics, i.e., the ER, MED, MRED, and NMED, but consumes more area, power,512and energy. The HOERAA shows the best accuracy performance in the metrics except the ER among513the LOA and its variants. The MED and NMED of the proposed adder are comparable to those of the514ECPETA, and the MRED of the proposed adder is the same as that of the LOA, OLOCA, HOERAA,515and CPETA. Importantly, our proposed adder outperforms the other approximate adders in terms of516all the accuracy metrics, the exceptions being its higher MED, MRED, and NMED than those of the517ECPETA. Specifically, the proposed HERLOA shows 1.31, 2.09, and 2.24 times better MED than the518LOA, ETAI, and LOAWA, respectively, and 1.16 times better MRED than the LOAWA and ETAI.519Table 1. Summary of performances of various 16-bit adders with n = 16 and k = 8.520 521Adder522 523Area524(µm2)525 526Delay527(ns)528 529Power530(µW)531 532PDP533(fJ)534 535EDP536(fJ·s)537 538ER539(%)540 541MED542 543MRED544 545NMED546 547RCA548 549157.8550 5512.23552 55345.0554 555100.4556 5572.24 × 10−7558 559N/A560 561N/A562 563N/A564 565N/A566 56759.4568 5696.06 × 10570 571−8572 573CLA574 575227.9576 5771.02578 57958.2580 581N/A582 583N/A584 585N/A586 587N/A588 58924.4590 59127.8592 5933.17 × 10594 595−8596 597LOA598 59997.0600 6011.14602 60389.98604 605110.9606 6074.38608 6091.69 × 10−3610 611LOAWA612 61388.6614 6151.09616 61722.0618 61924.0620 6212.62 × 10−8622 62389.98624 625189.6626 6275.08628 6292.89 × 10−3630 631OLOCA632 63385.4634 6351.14636 63722.9638 63926.1640 6412.98 × 10−8642 64399.13644 645115.0646 6474.38648 6491.75 × 10−3650 651HOERAA652 65389.3654 6551.16656 65723.8658 65927.6660 6613.20 × 10−8662 66398.84664 66595.01666 6674.38668 6691.45 × 10−3670 671ETAI672 673108.5674 6751.09676 67726.3678 67928.7680 6813.13 × 10−8682 68389.98684 685177.1686 6875.08688 6892.70 × 10−3690 691CPETA692 693113.9694 6951.14696 69728.0698 69931.9700 7013.64 × 10−8702 70386.65704 70588.5706 7074.38708 7091.35 × 10−3710 711ECPETA712 713116.5714 7151.26716 71729.7718 71937.4720 7214.71 × 10−8722 72386.65724 72581.5726 7273.70728 7291.24 × 10−3730 731HERLOA732 733104.6734 7351.14736 73726.6738 73930.3740 7413.45 × 10−8742 74384.43744 74584.8746 7474.38748 7491.29 × 10−3750 7514.2. Accuracy Analysis752For evaluating the accuracy of the proposed HERLOA in comparison with those of the other753seven approximate adders, we varied the design parameter k from 6 to 12 in order to alter the bitwidth of the precise part of the 16-bit adders and to extract the value of the ER, MED, MRED, and754NMED metrics.755Figure 6 shows the ERs of the approximate adders at various values of k. Clearly, the ER756decreases as the precise adder size k increases. Regardless of k, the LOA, LOAWA, and ETAI have an757identical ER, as do the CPETA and ECPETA, where the later ER is lower than that of the LOA. The758proposed HERLOA has the lowest ER and the OLOCA has the highest ER. The HOERAA shows759slightly better ER performance than the OLOCA and has the second highest ER. Specifically, the ERs760of the OLOCA and HOERAA reach 85% and 81%, respectively, and that of our HERLOA decreases761to 50% thanks to the proposed hybrid error reduction scheme at k = 12. The LOA and CPETA have762ERs of 68% and 57%, respectively, at the same k. We also plotted the line of Equation (2) in Figure 6763to determine the accuracy of the ER of our HERLOA as derived by this equation. The line is in very764good agreement with the simulated ERs at the various values of k.765 766Electronics 2020, 9, 471767 7689 of 13769 770100771Error Rate (%)772 7739077480775707766077750778407796780 7817782 7838784 7859786 78710788 78911790 79112792 793k794LOA795 796LOAWA797 798OLOCA799 800HOERAA801 802CPETA803 804ECPETA805 806HERLOA807 808Equation (2)809 810ETAI811 812Figure 6. Comparison of error rates of 16-bit approximate adders at various values of design813parameter k.814 815As reported in Table 1, the proposed adder has the second-best MED performance at k = 8. To816effectively demonstrate the MED of the proposed HERLOA and compare it with those of the other817adders, we plotted the improvements in the MED of the proposed adder in comparison with those of818the other seven approximate adders at various values of k in Figure 7. The MED improvement against819all the adders except the ECPETA increases with increasing k. Although the OLOCA has the highest820ER, its MED is similar to that of the LOA. Similarly, the HOERAA shows better MED performance821than the LOA, LOAWA, and ETAI in spite of worse ER performance than those adders. Moreover, in822terms of the MED, the HOERAA outperforms the OLOCA. The MED of our design is 30–43% better823than those of the LOA and OLOCA. Furthermore, at all k values, the MED improvement of our adder824reaches over 100% compared with the LOAWA and ETAI. Unfortunately, the proposed design shows8253–5% less MED performance against the ECPETA. The MED of the proposed design is comparable to826those the ETA variants, which include their own carry prediction schemes; the MED difference in this827case is ±5% at all k values.828 829MED Improvment (%)830 831140832120833100834808356083640837208380839-20840 8416842 8437844 8458846 8479848k849 85010851 852Proposed vs LOA853 854Proposed vs LOAWA855 856Proposed vs OLOCA857 858Proposed vs ETAI859 860Proposed vs CPETA861 862Proposed vs ECPETA863 86411865 86612867 868Proposed vs HOERAA869 870Figure 7. Improvements in the mean error distance (MED) of the proposed HERLOA in comparison871with those of seven approximate adders at various values of k.872 873Table 2 lists the MREDs of the seven approximate adders, including our proposed adder. For all874the adders, the MRED decreases as k increases. The MREDs of the LOA, OLOCA, HOERAA, CPETA,875 876Electronics 2020, 9, 471877 87810 of 13879 880and the proposed HERLOA are almost the same, whereas those of the LOAWA and ETAI are881relatively higher and that of the ECPETA is lower. Interestingly, the MREDs of the adders are almost882identical when they have the same carry prediction scheme. For example, the LOA, OLOCA,883HOERAA, and HERLOA include the AND-operation-based carry prediction scheme and have the884same MRED, and the LOAWA and ETAI do not include any carry prediction scheme (i.e., Cin = 0) and885have identical MREDs. The ECPETA includes the most accurate carry prediction scheme and shows886the best MRED performance among all the adders considered in this paper. The MRED difference887between the proposed adder and the ETAI/LOAWA remains unchanged as k increases. Specifically,888this MRED difference is are approximately 0.7 over the entire considered range of k values. Since the889MRED decreases as k increases, the percentage of MRED improvement increases as k increases. For890example, our design achieves MRED reductions of 12.3%, 16.0%, 25.3%, and 51.0%, at k values of 6,8918, 10, and 12, respectively.892Table 2. Mean relative error distances (MREDs) of 16-bit approximate adders at various values of893design parameter k.894 895k896 897LOA898 899LOAWA900 901OLOCA902 903HOERAA904 905ETAI906 907CPETA908 909EPCETA910 911HERLOA912 9136914 9155.77916 9176.47918 9195.77920 9215.76922 9236.47924 9255.76926 9275.06928 9295.76930 9317932 9335.12934 9355.81936 9375.12938 9395.12940 9415.81942 9435.12944 9454.42946 9475.12948 9498950 9514.38952 9535.08954 9554.38956 9574.38958 9595.08960 9614.38962 9633.70964 9654.38966 9679968 9693.81970 9714.53972 9733.81974 9753.81976 9774.53978 9793.81980 9813.05982 9833.81984 98510986 9872.89988 9893.62990 9912.89992 9932.89994 9953.62996 9972.89998 9992.061000 10012.891002 1003111004 10052.051006 10072.781008 10092.051010 10112.051012 10132.781014 10152.051016 10171.331018 10192.051020 1021121022 10231.411024 10252.131026 10271.411028 10291.411030 10312.131032 10331.411034 10350.621036 10371.411038 10394.3. Joint Analysis of Performance and Accuracy1040Power-NMED and energy-NMED products were introduced in [32] and [19], respectively, to1041assess the tradeoff among the power, energy, and accuracy of approximate adders. Here, we can take1042into account a new metric, the EDP-NMED product, to jointly analyze the tradeoff among energy,1043delay, and accuracy.1044Figure 8 shows the power-NMED, energy-NMED, and EDP-NMED products of the seven1045approximate adders with n = 16 and k = 8. All three products are normalized using corresponding1046values of the LOA to effectively demonstrate the tradeoffs. Impressively, the proposed HERLOA1047shows the best tradeoff performance, whereas the ETAI has the largest values of all three products.1048Specifically, the power-, energy-, and EDP-NMED products of the ETAI are 72%, 65%, and 58%,1049respectively, larger than those of the LOA. In contrast, all three products of the proposed design are105016% smaller than those of the LOA. The HOERAA, which has slightly less energy- and EDP-NMED1051products than the proposed adder, also shows a good tradeoff and is comparable to the proposed1052HERLOA. Although the ECPETA demonstrates the best NMED performance, the additional delay1053and power consumption of this adder that originate from its carry prediction scheme prevent it from1054having the best tradeoff metrics. As an example, the EDP-NMED product of the ECPETA is 30%1055larger than that of the proposed design. Specifically, the power-, energy-, and EDP-NMED products1056of the proposed approximate adder are, respectively, 2.06, 1.97, and 1.88 times better than those of1057the ETAI. Clearly, the excellent tradeoff between hardware cost and computation accuracy makes1058our adder design the most competitive among all the considered approximate adders, which have1059similar hardware architectures.1060 1061Electronics 2020, 9, 4711062 106311 of 131064 1065210661.721067 10681.651069 1070Normalized Value1071 10721.551073 10741.581075 10761.481077 10781.51079 10801.411081 10821.0910831.001084 108511086 10870.9810880.841089 10900.92 0.891091 10921.001093 10940.9810950.851096 10970.841098 10990.921100 11011.001102 11030.9911040.841105 11060.9811070.871108 11090.921110 11110.841112 11130.51114 111501116Power-NMED1117Product1118LOA1119 1120LOAWA1121 1122OLOCA1123 1124Energy-NMED1125Product1126HOERAA1127 1128ETAI1129 1130EDP-NMED1131Product1132CPETA1133 1134ECPETA1135 1136HERLOA1137 1138Figure 8. Normalized power-normalized mean error distance (NMED), energy-NMED, and energydelay product-NMED (EDP-NMED) products of approximate adders with n = 16 and k = 8.1139 11405. Conclusion1141In this paper, we have developed an accuracy enhanced lower-part OR adder with a hybrid error1142reduction scheme (termed the HERLOA) to significantly reduce the computation error while1143maintaining power and energy efficiencies. The proposed HERLOA replaces the OR gate with the1144XOR gate in the MSB of the approximate part and leverages two MSB inputs of the approximate part1145to decrease the approximation errors at the cost of a few digital gates. The proposed design is1146implemented in 65-nm technology to evaluate its performance; it is found to be 3, 2, and 2 times more1147energy-efficient, area-efficient, and power-efficient, respectively, than the RCA and CLA. In terms of1148accuracy, the proposed HERLOA outperforms the original LOA and the ETAI. Specifically, at a given1149design parameter value of k = 12, the ER of our adder decreases to 50%, whereas those of the other1150adders reach 68%, and at all the considered values of k, the MED improvement of our adder is more1151than 100% compared with the ETAI and LOAWA. Most importantly, the proposed adder shows an1152excellent design tradeoff between hardware cost and computation accuracy, as a result of which its1153power-NMED, energy-NMED, and EDP-NMED products are 2.06, 1.97, and 1.88 times better,1154respectively, than those of the ETAI. To sum up, our proposed adder outperforms all the approximate1155adders in a joint analysis of power, energy, EDP, and computation accuracy.1156Consequently, the proposed approximate adder with the novel hybrid error reduction scheme1157is found to be highly power- and energy-efficient while also having good computation accuracy.1158Therefore, our design is highly suitable for application to inherently error-resilient energy-efficient1159computing, such as DSP, deep learning, and neuromorphic computing [2–4,34,35].1160Author Contributions: Conceptualization, Y.K. and Y.S.Y.; methodology, H.S.; software, H.S. and Y.K.;1161validation, H.S.; formal analysis, Y.K.; investigation, Y.S.Y.; writing—original draft preparation, Y.K.; writing—1162review and editing, H.S. and Y.S.Y.; visualization, H.S.; supervision, Y.K.; funding acquisition, Y.K. All authors1163have read and agreed to the published version of the manuscript.1164Funding: This research was supported in part by Basic Science Research Program through the National Research1165Foundation of Korea (NRF) funded by the Ministry of Education (NRF-2019R1I1A3A01061266) and in part by1166the BK21 Plus Project (SW Human Resource Development Program for Supporting Smart Life) funded by the1167Ministry of Education, School of Computer Science and Engineering, Kyungpook National University, Korea1168(21A20131600005).1169Conflicts of Interest: The authors declare no conflict of interest.1170 1171Electronics 2020, 9, 4711172 117312 of 131174 1175References11761.11772.11783.11794.11805.11816.11827.11838.1184 11859.1186 118710.1188 118911.1190 119112.119213.119314.1194 119515.1196 119716.119817.1199 120018.