Team Ai
Datasetpublic

codekingpro/portable-devtools

sourceHugging Faceupdated 5mo agoView on Hugging Face
1likes14kdownloads
pipeline586 linesDownload Raw Back to cuda
1/*2 * SPDX-FileCopyrightText: Copyright (c) 2021 NVIDIA CORPORATION & AFFILIATES. All rights reserved.3 *4 * NVIDIA SOFTWARE LICENSE5 *6 * This license is a legal agreement between you and NVIDIA Corporation ("NVIDIA") and governs your use of the NVIDIA/CUDA C++ Library software and materials provided hereunder (“SOFTWARE”).7 *8 * This license can be accepted only by an adult of legal age of majority in the country in which the SOFTWARE is used. If you are under the legal age of majority, you must ask your parent or legal guardian to consent to this license. By taking delivery of the SOFTWARE, you affirm that you have reached the legal age of majority, you accept the terms of this license, and you take legal and financial responsibility for the actions of your permitted users.9 *10 * You agree to use the SOFTWARE only for purposes that are permitted by (a) this license, and (b) any applicable law, regulation or generally accepted practices or guidelines in the relevant jurisdictions.11 *12 * 1. LICENSE. Subject to the terms of this license, NVIDIA grants you a non-exclusive limited license to: (a) install and use the SOFTWARE, and (b) distribute the SOFTWARE subject to the distribution requirements described in this license. NVIDIA reserves all rights, title and interest in and to the SOFTWARE not expressly granted to you under this license.13 *14 * 2. DISTRIBUTION REQUIREMENTS. These are the distribution requirements for you to exercise the distribution grant:15 * a.      The terms under which you distribute the SOFTWARE must be consistent with the terms of this license, including (without limitation) terms relating to the license grant and license restrictions and protection of NVIDIA’s intellectual property rights.16 * b.      You agree to notify NVIDIA in writing of any known or suspected distribution or use of the SOFTWARE not in compliance with the requirements of this license, and to enforce the terms of your agreements with respect to distributed SOFTWARE.17 *18 * 3. LIMITATIONS. Your license to use the SOFTWARE is restricted as follows:19 * a.      The SOFTWARE is licensed for you to develop applications only for use in systems with NVIDIA GPUs.20 * b.      You may not reverse engineer, decompile or disassemble, or remove copyright or other proprietary notices from any portion of the SOFTWARE or copies of the SOFTWARE.21 * c.      You may not modify or create derivative works of any portion of the SOFTWARE.22 * d.      You may not bypass, disable, or circumvent any technical measure, encryption, security, digital rights management or authentication mechanism in the SOFTWARE.23 * e.      You may not use the SOFTWARE in any manner that would cause it to become subject to an open source software license. As examples, licenses that require as a condition of use, modification, and/or distribution that the SOFTWARE be (i) disclosed or distributed in source code form; (ii) licensed for the purpose of making derivative works; or (iii) redistributable at no charge.24 * f.      Unless you have an agreement with NVIDIA for this purpose, you may not use the SOFTWARE with any system or application where the use or failure of the system or application can reasonably be expected to threaten or result in personal injury, death, or catastrophic loss. Examples include use in avionics, navigation, military, medical, life support or other life critical applications. NVIDIA does not design, test or manufacture the SOFTWARE for these critical uses and NVIDIA shall not be liable to you or any third party, in whole or in part, for any claims or damages arising from such uses.25 * g.      You agree to defend, indemnify and hold harmless NVIDIA and its affiliates, and their respective employees, contractors, agents, officers and directors, from and against any and all claims, damages, obligations, losses, liabilities, costs or debt, fines, restitutions and expenses (including but not limited to attorney’s fees and costs incident to establishing the right of indemnification) arising out of or related to use of the SOFTWARE outside of the scope of this Agreement, or not in compliance with its terms.26 *27 * 4. PRE-RELEASE. SOFTWARE versions identified as alpha, beta, preview, early access or otherwise as pre-release may not be fully functional, may contain errors or design flaws, and may have reduced or different security, privacy, availability, and reliability standards relative to commercial versions of NVIDIA software and materials. You may use a pre-release SOFTWARE version at your own risk, understanding that these versions are not intended for use in production or business-critical systems.28 *29 * 5. OWNERSHIP. The SOFTWARE and the related intellectual property rights therein are and will remain the sole and exclusive property of NVIDIA or its licensors. The SOFTWARE is copyrighted and protected by the laws of the United States and other countries, and international treaty provisions. NVIDIA may make changes to the SOFTWARE, at any time without notice, but is not obligated to support or update the SOFTWARE.30 *31 * 6. COMPONENTS UNDER OTHER LICENSES. The SOFTWARE may include NVIDIA or third-party components with separate legal notices or terms as may be described in proprietary notices accompanying the SOFTWARE. If and to the extent there is a conflict between the terms in this license and the license terms associated with a component, the license terms associated with the components control only to the extent necessary to resolve the conflict.32 *33 * 7. FEEDBACK. You may, but don’t have to, provide to NVIDIA any Feedback. “Feedback” means any suggestions, bug fixes, enhancements, modifications, feature requests or other feedback regarding the SOFTWARE. For any Feedback that you voluntarily provide, you hereby grant NVIDIA and its affiliates a perpetual, non-exclusive, worldwide, irrevocable license to use, reproduce, modify, license, sublicense (through multiple tiers of sublicensees), and distribute (through multiple tiers of distributors) the Feedback without the payment of any royalties or fees to you. NVIDIA will use Feedback at its choice.34 *35 * 8. NO WARRANTIES. THE SOFTWARE IS PROVIDED "AS IS" WITHOUT ANY EXPRESS OR IMPLIED WARRANTY OF ANY KIND INCLUDING, BUT NOT LIMITED TO, WARRANTIES OF MERCHANTABILITY, NONINFRINGEMENT, OR FITNESS FOR A PARTICULAR PURPOSE. NVIDIA DOES NOT WARRANT THAT THE SOFTWARE WILL MEET YOUR REQUIREMENTS OR THAT THE OPERATION THEREOF WILL BE UNINTERRUPTED OR ERROR-FREE, OR THAT ALL ERRORS WILL BE CORRECTED.36 *37 * 9. LIMITATIONS OF LIABILITY. TO THE MAXIMUM EXTENT PERMITTED BY LAW, NVIDIA AND ITS AFFILIATES SHALL NOT BE LIABLE FOR ANY SPECIAL, INCIDENTAL, PUNITIVE OR CONSEQUENTIAL DAMAGES, OR ANY LOST PROFITS, PROJECT DELAYS, LOSS OF USE, LOSS OF DATA OR LOSS OF GOODWILL, OR THE COSTS OF PROCURING SUBSTITUTE PRODUCTS, ARISING OUT OF OR IN CONNECTION WITH THIS LICENSE OR THE USE OR PERFORMANCE OF THE SOFTWARE, WHETHER SUCH LIABILITY ARISES FROM ANY CLAIM BASED UPON BREACH OF CONTRACT, BREACH OF WARRANTY, TORT (INCLUDING NEGLIGENCE), PRODUCT LIABILITY OR ANY OTHER CAUSE OF ACTION OR THEORY OF LIABILITY, EVEN IF NVIDIA HAS PREVIOUSLY BEEN ADVISED OF, OR COULD REASONABLY HAVE FORESEEN, THE POSSIBILITY OF SUCH DAMAGES. IN NO EVENT WILL NVIDIA’S AND ITS AFFILIATES TOTAL CUMULATIVE LIABILITY UNDER OR ARISING OUT OF THIS LICENSE EXCEED US$10.00. THE NATURE OF THE LIABILITY OR THE NUMBER OF CLAIMS OR SUITS SHALL NOT ENLARGE OR EXTEND THIS LIMIT.38 *39 * 10. TERMINATION. Your rights under this license will terminate automatically without notice from NVIDIA if you fail to comply with any term and condition of this license or if you commence or participate in any legal proceeding against NVIDIA with respect to the SOFTWARE. NVIDIA may terminate this license with advance written notice to you if NVIDIA decides to no longer provide the SOFTWARE in a country or, in NVIDIA’s sole discretion, the continued use of it is no longer commercially viable. Upon any termination of this license, you agree to promptly discontinue use of the SOFTWARE and destroy all copies in your possession or control. Your prior distributions in accordance with this license are not affected by the termination of this license. All provisions of this license will survive termination, except for the license granted to you.40 *41 * 11. APPLICABLE LAW. This license will be governed in all respects by the laws of the United States and of the State of Delaware as those laws are applied to contracts entered into and performed entirely within Delaware by Delaware residents, without regard to the conflicts of laws principles. The United Nations Convention on Contracts for the International Sale of Goods is specifically disclaimed. You agree to all terms of this Agreement in the English language. The state or federal courts residing in Santa Clara County, California shall have exclusive jurisdiction over any dispute or claim arising out of this license. Notwithstanding this, you agree that NVIDIA shall still be allowed to apply for injunctive remedies or an equivalent type of urgent legal relief in any jurisdiction.42 *43 * 12. NO ASSIGNMENT. This license and your rights and obligations thereunder may not be assigned by you by any means or operation of law without NVIDIA’s permission. Any attempted assignment not approved by NVIDIA in writing shall be void and of no effect.44 *45 * 13. EXPORT. The SOFTWARE is subject to United States export laws and regulations. You agree that you will not ship, transfer or export the SOFTWARE into any country, or use the SOFTWARE in any manner, prohibited by the United States Bureau of Industry and Security or economic sanctions regulations administered by the U.S. Department of Treasury’s Office of Foreign Assets Control (OFAC), or any applicable export laws, restrictions or regulations. These laws include restrictions on destinations, end users and end use. By accepting this license, you confirm that you are not a resident or citizen of any country currently embargoed by the U.S. and that you are not otherwise prohibited from receiving the SOFTWARE.46 *47 * 14. GOVERNMENT USE. The SOFTWARE has been developed entirely at private expense and is “commercial items” consisting of “commercial computer software” and “commercial computer software documentation” provided with RESTRICTED RIGHTS. Use, duplication or disclosure by the U.S. Government or a U.S. Government subcontractor is subject to the restrictions in this license pursuant to DFARS 227.7202-3(a) or as set forth in subparagraphs (b)(1) and (2) of the Commercial Computer Software - Restricted Rights clause at FAR 52.227-19, as applicable. Contractor/manufacturer is NVIDIA, 2788 San Tomas Expressway, Santa Clara, CA 95051.48 *49 * 15. ENTIRE AGREEMENT. This license is the final, complete and exclusive agreement between the parties relating to the subject matter of this license and supersedes all prior or contemporaneous understandings and agreements relating to this subject matter, whether oral or written. If any court of competent jurisdiction determines that any provision of this license is illegal, invalid or unenforceable, the remaining provisions will remain in full force and effect. This license may only be modified in a writing signed by an authorized representative of each party.50 *51 * (v. August 20, 2021)52 */53#ifndef _CUDA_PIPELINE54#define _CUDA_PIPELINE55 56#include "barrier"57#include "atomic"58#include "std/chrono"59 60_LIBCUDACXX_BEGIN_NAMESPACE_CUDA61 62    // Forward declaration in barrier of pipeline63    enum class pipeline_role {64        producer,65        consumer66    };67 68    template<thread_scope _Scope>69    struct __pipeline_stage {70        barrier<_Scope> __produced;71        barrier<_Scope> __consumed;72    };73 74    template<thread_scope _Scope, uint8_t _Stages_count>75    class pipeline_shared_state {76    public:77        pipeline_shared_state() = default;78        pipeline_shared_state(const pipeline_shared_state &) = delete;79        pipeline_shared_state(pipeline_shared_state &&) = delete;80        pipeline_shared_state & operator=(pipeline_shared_state &&) = delete;81        pipeline_shared_state & operator=(const pipeline_shared_state &) =  delete;82 83    private:84        __pipeline_stage<_Scope> __stages[_Stages_count];85        atomic<uint32_t, _Scope> __refcount;86 87        template<thread_scope _Pipeline_scope>88        friend class pipeline;89 90        template<class _Group, thread_scope _Pipeline_scope, uint8_t _Pipeline_stages_count>91        friend _LIBCUDACXX_INLINE_VISIBILITY92        pipeline<_Pipeline_scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Pipeline_scope, _Pipeline_stages_count> * __shared_state);93 94        template<class _Group, thread_scope _Pipeline_scope, uint8_t _Pipeline_stages_count>95        friend _LIBCUDACXX_INLINE_VISIBILITY96        pipeline<_Pipeline_scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Pipeline_scope, _Pipeline_stages_count> * __shared_state, size_t __producer_count);97 98        template<class _Group, thread_scope _Pipeline_scope, uint8_t _Pipeline_stages_count>99        friend _LIBCUDACXX_INLINE_VISIBILITY100        pipeline<_Pipeline_scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Pipeline_scope, _Pipeline_stages_count> * __shared_state, pipeline_role __role);101    };102 103    struct __pipeline_asm_helper {104        _LIBCUDACXX_DEVICE105        static inline uint32_t __lane_id()106        {107            NV_IF_ELSE_TARGET(108                NV_IS_DEVICE,109                (110                    uint32_t __lane_id;111                    asm volatile ("mov.u32 %0, %%laneid;" : "=r"(__lane_id));112                    return __lane_id;113                ),114                (115                    return 0;116                )117            )118        }119    };120 121    template<thread_scope _Scope>122    class pipeline {123    public:124        pipeline(pipeline &&) = default;125        pipeline(const pipeline &) = delete;126        pipeline & operator=(pipeline &&) = delete;127        pipeline & operator=(const pipeline &) = delete;128 129        _LIBCUDACXX_INLINE_VISIBILITY130        ~pipeline()131        {132            if (__active) {133                (void)quit();134            }135        }136 137        _LIBCUDACXX_INLINE_VISIBILITY138        bool quit()139        {140            bool __elected;141            uint32_t __sub_count;142NV_IF_TARGET(NV_IS_DEVICE,143            const uint32_t __match_mask = __match_any_sync(__activemask(), reinterpret_cast<uintptr_t>(__shared_state_get_refcount()));144            const uint32_t __elected_id = __ffs(__match_mask) - 1;145            __elected = (__pipeline_asm_helper::__lane_id() == __elected_id);146            __sub_count = __popc(__match_mask);147,148            __elected = true;149            __sub_count = 1;150)151            bool __released = false;152            if (__elected) {153                const uint32_t __old = __shared_state_get_refcount()->fetch_sub(__sub_count);154                const bool __last = (__old == __sub_count);155                if (__last) {156                    for (uint8_t __stage = 0; __stage < __stages_count; ++__stage) {157                        __shared_state_get_stage(__stage)->__produced.~barrier();158                        __shared_state_get_stage(__stage)->__consumed.~barrier();159                    }160                    __released = true;161                }162            }163            __active = false;164            return __released;165        }166 167        _LIBCUDACXX_INLINE_VISIBILITY168        void producer_acquire()169        {170            barrier<_Scope> & __stage_barrier = __shared_state_get_stage(__head)->__consumed;171            __stage_barrier.wait_parity(__consumed_phase_parity);172        }173 174        _LIBCUDACXX_INLINE_VISIBILITY175        void producer_commit()176        {177            barrier<_Scope> & __stage_barrier = __shared_state_get_stage(__head)->__produced;178            (void)__memcpy_completion_impl::__defer(__completion_mechanism::__async_group, __single_thread_group{}, 0, __stage_barrier);179            (void)__stage_barrier.arrive();180            if (++__head == __stages_count) {181                __head = 0;182                __consumed_phase_parity = !__consumed_phase_parity;183            }184        }185 186        _LIBCUDACXX_INLINE_VISIBILITY187        void consumer_wait()188        {189            barrier<_Scope> & __stage_barrier = __shared_state_get_stage(__tail)->__produced;190            __stage_barrier.wait_parity(__produced_phase_parity);191        }192 193        _LIBCUDACXX_INLINE_VISIBILITY194        void consumer_release()195        {196            (void)__shared_state_get_stage(__tail)->__consumed.arrive();197            if (++__tail == __stages_count) {198                __tail = 0;199                __produced_phase_parity = !__produced_phase_parity;200            }201        }202 203        template<class _Rep, class _Period>204        _LIBCUDACXX_INLINE_VISIBILITY205        bool consumer_wait_for(const _CUDA_VSTD::chrono::duration<_Rep, _Period> & __duration)206        {207            barrier<_Scope> & __stage_barrier = __shared_state_get_stage(__tail)->__produced;208            return _CUDA_VSTD::__libcpp_thread_poll_with_backoff(209                        _CUDA_VSTD::__barrier_poll_tester_parity<barrier<_Scope>>(210                            &__stage_barrier,211                            __produced_phase_parity),212                        _CUDA_VSTD::chrono::duration_cast<_CUDA_VSTD::chrono::nanoseconds>(__duration)213            );214        }215 216        template<class _Clock, class _Duration>217        _LIBCUDACXX_INLINE_VISIBILITY218        bool consumer_wait_until(const _CUDA_VSTD::chrono::time_point<_Clock, _Duration> & __time_point)219        {220            return consumer_wait_for(__time_point - _Clock::now());221        }222 223    private:224        uint8_t __head               : 8;225        uint8_t __tail               : 8;226        const uint8_t __stages_count : 8;227        bool __consumed_phase_parity : 1;228        bool __produced_phase_parity : 1;229        bool __active                : 1;230        // TODO: Remove partitioned on next ABI break231        const bool __partitioned     : 1;232        char * const __shared_state;233 234 235        _LIBCUDACXX_INLINE_VISIBILITY236        pipeline(char * __shared_state, uint8_t __stages_count, bool __partitioned)237            : __head(0)238            , __tail(0)239            , __stages_count(__stages_count)240            , __consumed_phase_parity(true)241            , __produced_phase_parity(false)242            , __active(true)243            , __partitioned(__partitioned)244            , __shared_state(__shared_state)245        {}246 247        _LIBCUDACXX_INLINE_VISIBILITY248        __pipeline_stage<_Scope> * __shared_state_get_stage(uint8_t __stage)249        {250            ptrdiff_t __stage_offset = __stage * sizeof(__pipeline_stage<_Scope>);251            return reinterpret_cast<__pipeline_stage<_Scope>*>(__shared_state + __stage_offset);252        }253 254        _LIBCUDACXX_INLINE_VISIBILITY255        atomic<uint32_t, _Scope> * __shared_state_get_refcount()256        {257            ptrdiff_t __refcount_offset = __stages_count * sizeof(__pipeline_stage<_Scope>);258            return reinterpret_cast<atomic<uint32_t, _Scope>*>(__shared_state + __refcount_offset);259        }260 261        template<class _Group, thread_scope _Pipeline_scope, uint8_t _Pipeline_stages_count>262        friend _LIBCUDACXX_INLINE_VISIBILITY263        pipeline<_Pipeline_scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Pipeline_scope, _Pipeline_stages_count> * __shared_state);264 265        template<class _Group, thread_scope _Pipeline_scope, uint8_t _Pipeline_stages_count>266        friend _LIBCUDACXX_INLINE_VISIBILITY267        pipeline<_Pipeline_scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Pipeline_scope, _Pipeline_stages_count> * __shared_state, size_t __producer_count);268 269        template<class _Group, thread_scope _Pipeline_scope, uint8_t _Pipeline_stages_count>270        friend _LIBCUDACXX_INLINE_VISIBILITY271        pipeline<_Pipeline_scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Pipeline_scope, _Pipeline_stages_count> * __shared_state, pipeline_role __role);272    };273 274    template<class _Group, thread_scope _Scope, uint8_t _Stages_count>275    _LIBCUDACXX_INLINE_VISIBILITY276    pipeline<_Scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Scope, _Stages_count> * __shared_state)277    {278        const uint32_t __group_size = static_cast<uint32_t>(__group.size());279        const uint32_t __thread_rank = static_cast<uint32_t>(__group.thread_rank());280 281        if (__thread_rank == 0) {282            for (uint8_t __stage = 0; __stage < _Stages_count; ++__stage) {283                init(&__shared_state->__stages[__stage].__consumed, __group_size);284                init(&__shared_state->__stages[__stage].__produced, __group_size);285            }286            __shared_state->__refcount.store(__group_size, std::memory_order_relaxed);287        }288        __group.sync();289 290        return pipeline<_Scope>(reinterpret_cast<char*>(__shared_state->__stages), _Stages_count, false);291    }292 293    template<class _Group, thread_scope _Scope, uint8_t _Stages_count>294    _LIBCUDACXX_INLINE_VISIBILITY295    pipeline<_Scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Scope, _Stages_count> * __shared_state, size_t __producer_count)296    {297        const uint32_t __group_size = static_cast<uint32_t>(__group.size());298        const uint32_t __thread_rank = static_cast<uint32_t>(__group.thread_rank());299 300        if (__thread_rank == 0) {301            const size_t __consumer_count = __group_size - __producer_count;302            for (uint8_t __stage = 0; __stage < _Stages_count; ++__stage) {303                init(&__shared_state->__stages[__stage].__consumed, __consumer_count);304                init(&__shared_state->__stages[__stage].__produced, __producer_count);305            }306            __shared_state->__refcount.store(__group_size, std::memory_order_relaxed);307        }308        __group.sync();309 310        return pipeline<_Scope>(reinterpret_cast<char*>(__shared_state->__stages), _Stages_count, true);311    }312 313    template<class _Group, thread_scope _Scope, uint8_t _Stages_count>314    _LIBCUDACXX_INLINE_VISIBILITY315    pipeline<_Scope> make_pipeline(const _Group & __group, pipeline_shared_state<_Scope, _Stages_count> * __shared_state, pipeline_role __role)316    {317        const uint32_t __group_size = static_cast<uint32_t>(__group.size());318        const uint32_t __thread_rank = static_cast<uint32_t>(__group.thread_rank());319 320        if (__thread_rank == 0) {321            __shared_state->__refcount.store(0, std::memory_order_relaxed);322        }323        __group.sync();324 325        if (__role == pipeline_role::producer) {326            bool __elected;327            uint32_t __add_count;328NV_IF_TARGET(NV_IS_DEVICE,329            const uint32_t __match_mask = __match_any_sync(__activemask(), reinterpret_cast<uintptr_t>(&__shared_state->__refcount));330            const uint32_t __elected_id = __ffs(__match_mask) - 1;331            __elected = (__pipeline_asm_helper::__lane_id() == __elected_id);332            __add_count = __popc(__match_mask);333,334            __elected = true;335            __add_count = 1;336)337            if (__elected) {338                (void)__shared_state->__refcount.fetch_add(__add_count, std::memory_order_relaxed);339            }340        }341        __group.sync();342 343        if (__thread_rank == 0) {344            const uint32_t __producer_count = __shared_state->__refcount.load(std::memory_order_relaxed);345            const uint32_t __consumer_count = __group_size - __producer_count;346            for (uint8_t __stage = 0; __stage < _Stages_count; ++__stage) {347                init(&__shared_state->__stages[__stage].__consumed, __consumer_count);348                init(&__shared_state->__stages[__stage].__produced, __producer_count);349            }350            __shared_state->__refcount.store(__group_size, std::memory_order_relaxed);351        }352        __group.sync();353 354        return pipeline<_Scope>(reinterpret_cast<char*>(__shared_state->__stages), _Stages_count, true);355    }356 357_LIBCUDACXX_END_NAMESPACE_CUDA358 359_LIBCUDACXX_BEGIN_NAMESPACE_CUDA_DEVICE360 361    template<uint8_t _Prior>362    _LIBCUDACXX_DEVICE363    void __pipeline_consumer_wait(pipeline<thread_scope_thread> & __pipeline);364 365    _LIBCUDACXX_DEVICE366    inline void __pipeline_consumer_wait(pipeline<thread_scope_thread> & __pipeline, uint8_t __prior);367 368_LIBCUDACXX_END_NAMESPACE_CUDA_DEVICE369 370_LIBCUDACXX_BEGIN_NAMESPACE_CUDA371 372    template<>373    class pipeline<thread_scope_thread> {374    public:375        pipeline(pipeline &&) = default;376        pipeline(const pipeline &) = delete;377        pipeline & operator=(pipeline &&) = delete;378        pipeline & operator=(const pipeline &) = delete;379 380        _LIBCUDACXX_INLINE_VISIBILITY381        ~pipeline() {}382 383        _LIBCUDACXX_INLINE_VISIBILITY384        bool quit()385        {386            return true;387        }388 389        _LIBCUDACXX_INLINE_VISIBILITY390        void producer_acquire() {}391 392        _LIBCUDACXX_INLINE_VISIBILITY393        void producer_commit()394        {395NV_IF_TARGET(NV_PROVIDES_SM_80,396            asm volatile ("cp.async.commit_group;");397            ++__head;398)399        }400 401        _LIBCUDACXX_INLINE_VISIBILITY402        void consumer_wait()403        {404NV_IF_TARGET(NV_PROVIDES_SM_80,405            if (__head == __tail) {406                return;407            }408 409            const uint8_t __prior = __head - __tail - 1;410            device::__pipeline_consumer_wait(*this, __prior);411            ++__tail;412)413        }414 415        _LIBCUDACXX_INLINE_VISIBILITY416        void consumer_release() {}417 418        template<class _Rep, class _Period>419        _LIBCUDACXX_INLINE_VISIBILITY420        bool consumer_wait_for(const _CUDA_VSTD::chrono::duration<_Rep, _Period> & __duration)421        {422            (void)__duration;423            consumer_wait();424            return true;425        }426 427        template<class _Clock, class _Duration>428        _LIBCUDACXX_INLINE_VISIBILITY429        bool consumer_wait_until(const _CUDA_VSTD::chrono::time_point<_Clock, _Duration> & __time_point)430        {431            (void)__time_point;432            consumer_wait();433            return true;434        }435 436    private:437        uint8_t __head;438        uint8_t __tail;439 440        _LIBCUDACXX_INLINE_VISIBILITY441        pipeline()442            : __head(0)443            , __tail(0)444        {}445 446        friend _LIBCUDACXX_INLINE_VISIBILITY inline pipeline<thread_scope_thread> make_pipeline();447 448        template<uint8_t _Prior>449        friend _LIBCUDACXX_INLINE_VISIBILITY450        void pipeline_consumer_wait_prior(pipeline<thread_scope_thread> & __pipeline);451 452        template<class _Group, thread_scope _Pipeline_scope, uint8_t _Pipeline_stages_count>453        friend _LIBCUDACXX_INLINE_VISIBILITY454        pipeline<_Pipeline_scope> __make_pipeline(const _Group & __group, pipeline_shared_state<_Pipeline_scope, _Pipeline_stages_count> * __shared_state);455    };456 457_LIBCUDACXX_END_NAMESPACE_CUDA458 459_LIBCUDACXX_BEGIN_NAMESPACE_CUDA_DEVICE460 461    template<uint8_t _Prior>462    _LIBCUDACXX_DEVICE463    void __pipeline_consumer_wait(pipeline<thread_scope_thread> & __pipeline)464    {465        (void)__pipeline;466NV_IF_TARGET(NV_PROVIDES_SM_80,467        constexpr uint8_t __max_prior = 8;468 469        asm volatile ("cp.async.wait_group %0;"470            :471            : "n"(_Prior < __max_prior ? _Prior : __max_prior));472)473    }474 475    _LIBCUDACXX_DEVICE476    inline void __pipeline_consumer_wait(pipeline<thread_scope_thread> & __pipeline, uint8_t __prior)477    {478        switch (__prior) {479        case 0:  device::__pipeline_consumer_wait<0>(__pipeline); break;480        case 1:  device::__pipeline_consumer_wait<1>(__pipeline); break;481        case 2:  device::__pipeline_consumer_wait<2>(__pipeline); break;482        case 3:  device::__pipeline_consumer_wait<3>(__pipeline); break;483        case 4:  device::__pipeline_consumer_wait<4>(__pipeline); break;484        case 5:  device::__pipeline_consumer_wait<5>(__pipeline); break;485        case 6:  device::__pipeline_consumer_wait<6>(__pipeline); break;486        case 7:  device::__pipeline_consumer_wait<7>(__pipeline); break;487        default: device::__pipeline_consumer_wait<8>(__pipeline); break;488        }489    }490 491_LIBCUDACXX_END_NAMESPACE_CUDA_DEVICE492 493_LIBCUDACXX_BEGIN_NAMESPACE_CUDA494 495    _LIBCUDACXX_INLINE_VISIBILITY496    inline pipeline<thread_scope_thread> make_pipeline()497    {498        return pipeline<thread_scope_thread>();499    }500 501    template<uint8_t _Prior>502    _LIBCUDACXX_INLINE_VISIBILITY503    void pipeline_consumer_wait_prior(pipeline<thread_scope_thread> & __pipeline)504    {505        NV_IF_TARGET(NV_PROVIDES_SM_80,506            device::__pipeline_consumer_wait<_Prior>(__pipeline);507            __pipeline.__tail = __pipeline.__head - _Prior;508        )509    }510 511    template<thread_scope _Scope>512    _LIBCUDACXX_INLINE_VISIBILITY513    void pipeline_producer_commit(pipeline<thread_scope_thread> & __pipeline, barrier<_Scope> & __barrier)514    {515        (void)__pipeline;516        NV_IF_TARGET(NV_PROVIDES_SM_80,(517            (void)__memcpy_completion_impl::__defer(__completion_mechanism::__async_group, __single_thread_group{}, 0, __barrier);518        ));519    }520 521    template<typename _Group, class _Tp, typename _Size, thread_scope _Scope>522    _LIBCUDACXX_INLINE_VISIBILITY523    async_contract_fulfillment __memcpy_async_pipeline(_Group const & __group, _Tp * __destination, _Tp const * __source, _Size __size, pipeline<_Scope> & __pipeline) {524        // 1. Set the completion mechanisms that can be used.525        //526        //    Do not (yet) allow async_bulk_group completion. Do not allow527        //    mbarrier_complete_tx completion, even though it may be possible if528        //    the pipeline has stage barriers in shared memory.529        _CUDA_VSTD::uint32_t __allowed_completions = _CUDA_VSTD::uint32_t(__completion_mechanism::__async_group);530 531        // Alignment: Use the maximum of the alignment of _Tp and that of a possible cuda::aligned_size_t.532        constexpr _CUDA_VSTD::size_t __size_align = __get_size_align<_Size>::align;533        constexpr _CUDA_VSTD::size_t __align = (alignof(_Tp) < __size_align) ? __size_align : alignof(_Tp);534        // Cast to char pointers. We don't need the type for alignment anymore and535        // erasing the types reduces the number of instantiations of down-stream536        // functions.537        char * __dest_char = reinterpret_cast<char*>(__destination);538        char const * __src_char = reinterpret_cast<char const *>(__source);539 540        // 2. Issue actual copy instructions.541        auto __cm =  __dispatch_memcpy_async<__align>(__group, __dest_char, __src_char, __size, __allowed_completions);542 543        // 3. No need to synchronize with copy instructions.544        return __memcpy_completion_impl::__defer(__cm, __group, __size, __pipeline);545    }546 547    template<typename _Group, class _Type, thread_scope _Scope>548    _LIBCUDACXX_INLINE_VISIBILITY549    async_contract_fulfillment memcpy_async(_Group const & __group, _Type * __destination, _Type const * __source, std::size_t __size, pipeline<_Scope> & __pipeline) {550        return __memcpy_async_pipeline(__group, __destination, __source, __size, __pipeline);551    }552 553    template<typename _Group, class _Type, std::size_t _Alignment, thread_scope _Scope, std::size_t _Larger_alignment = (alignof(_Type) > _Alignment) ? alignof(_Type) : _Alignment>554    _LIBCUDACXX_INLINE_VISIBILITY555    async_contract_fulfillment memcpy_async(_Group const & __group, _Type * __destination, _Type const * __source, aligned_size_t<_Alignment> __size, pipeline<_Scope> & __pipeline) {556        return __memcpy_async_pipeline(__group, __destination, __source, __size, __pipeline);557    }558 559    template<class _Type, typename _Size, thread_scope _Scope>560    _LIBCUDACXX_INLINE_VISIBILITY561    async_contract_fulfillment memcpy_async(_Type * __destination, _Type const * __source, _Size __size, pipeline<_Scope> & __pipeline) {562        return __memcpy_async_pipeline(__single_thread_group{}, __destination, __source, __size, __pipeline);563    }564 565    template<typename _Group, thread_scope _Scope>566    _LIBCUDACXX_INLINE_VISIBILITY567    async_contract_fulfillment memcpy_async(_Group const & __group, void * __destination, void const * __source, std::size_t __size, pipeline<_Scope> & __pipeline) {568        return __memcpy_async_pipeline(__group, reinterpret_cast<char *>(__destination), reinterpret_cast<char const *>(__source), __size, __pipeline);569    }570 571    template<typename _Group, std::size_t _Alignment, thread_scope _Scope>572    _LIBCUDACXX_INLINE_VISIBILITY573    async_contract_fulfillment memcpy_async(_Group const & __group, void * __destination, void const * __source, aligned_size_t<_Alignment> __size, pipeline<_Scope> & __pipeline) {574        return __memcpy_async_pipeline(__group, reinterpret_cast<char*>(__destination), reinterpret_cast<char const *>(__source), __size, __pipeline);575    }576 577    template<typename _Size, thread_scope _Scope>578    _LIBCUDACXX_INLINE_VISIBILITY579    async_contract_fulfillment memcpy_async(void * __destination, void const * __source, _Size __size, pipeline<_Scope> & __pipeline) {580        return __memcpy_async_pipeline(__single_thread_group{}, reinterpret_cast<char*>(__destination), reinterpret_cast<char const *>(__source), __size, __pipeline);581    }582 583_LIBCUDACXX_END_NAMESPACE_CUDA584 585#endif //_CUDA_PIPELINE586 
codekingpro/portable-devtools · Team Ai