# Terms of Service for Kimi OpenPlatform Source: https://platform.kimi.ai/docs/agreement/modeluse Review the rules for accounts, API usage, billing, data, compliance, and the rights and obligations of Kimi Open Platform users. **Last Updated: July 30th, 2026** Thank you for using Kimi OpenPlatform! The Terms of Service (**"the terms"**) constitute a legal agreement between Moonshot AI PTE. LTD. (**"we" or "Moonshot AI"**) and you (**"Customer"**) that governs the use of our APIs and other Services (**"the Services"**) by businesses and developers. The terms also reference and incorporate the "Kimi OpenPlatform Privacy Policy" and any additional guidelines or policies that we may provide in writing, and any ordering document signed by both you and us for the purchase of the Services (referred to as an "Order Form"). Together, these documents form the "**Agreement**." For more information on how we collect, use, and protect your personal information, you can read [Kimi OpenPlatform Privacy Policy](https://platform.kimi.ai/docs/agreement/userprivacy). By using the Services in any manner, especially placing an order, you agree that you have clearly read, understood the terms and agree to be bound by the terms of this Agreement. You further represent that you have the legal capacity to enter into contracts. If you are entering into the terms on behalf of an entity, you also represent that you have the legal authority to bind that entity. ## 1.Services We provide you with a non-exclusive license to access and utilize the Services for the duration of the term. This license allows you to use Moonshot AI's application programming interfaces (**"APIs"**) to integrate the Services into your own applications, products, or Services (each referred to as a **"Customer Application"**) and to offer those Customer Applications to **End Users**. The term "Services" encompasses any Services for businesses and developers that we make available for purchase or use, as well as any associated software, tools, developer Services, documentation, and websites, but does not include any Third Party Offerings.We possess all rights, titles, and interests in the Services. You are granted only the rights to use the Services as specifically provided in this Agreement. You must provide accurate and up-to-date account information. You are fully responsible for all activities under your account, including the activities of any End User or anyone else. You must not share your account credentials or make your account accessible to others,including not transferring, lending, leasing, or providing it to anyone for use in any form without authorization. If any loss occurs due to your failure to properly manage and safeguard the account, such as the account being stolen or used illegally, you will be held responsible for such loss. You must promptly notify us if you become aware of any unauthorized access to or use of your account or the Services. ## 2.Who can use the Services To use the Services as an individual, you must be at least 18 years old or the minimum age required by your country to consent to such use. If the account is used for enterprise, you must ensure that you have authorization from the business. In addition, you should not be the subject of any trade restrictions, sanctions, or other legal or regulatory restrictions imposed by any country, international organization, or region. ## 3.Usage policies In using the Services, you must comply with all applicable laws as well as our documentation, guidelines, or policies we make available to you. For example, you will not, and will not permit End Users to: **3.1 Use the Services for any illegal or improper purposes or activities that violate applicable laws and regulations, harm the legitimate rights or interests of us or anyone, including:** (1) For purposes that may have a harmful impact on physical or spiritual health or that violate ethical principles. (2) For promoting or engaging in any illegal activity, including terrorism, the exploitation or harm of children and the development or distribution of illegal substances, goods, or Services. (3) For engaging in activities that infringe upon intellectual property rights, unfair competitive practices like violations of trade secrets and business ethics. (4) For fraudulent, deceptive, or misleading activities. (5) To spam, bully, harass, defame, sexualize children, or promote violence, hatred or the suffering of others. (6) For compromising the privacy of others. (7) For any other uses prohibited or restricted by applicable laws and regulations, or that may harm legitimate interests of us or any other. **3.2 Engage in the following activities that endanger network security, operational security or create risks of the Services:** (1) Engaging in activities that endanger network security, such as unauthorized access to networks, interference with normal network functions, or theft of network data. (2) Providing programs or tools specifically designed for engaging in activities that endanger network security, such as network intrusion, interference with normal network functions and protective measures, or theft of network data. (3) Helping endanger network security with technical support, advertising promotion, payment settlement, or other assistance. (4) Reverse engineering, decompiling, disassembling, translating, or otherwise attempting to discover the source code, models, algorithms, or underlying components of this Services' system. (5) For developing, serving, or creating applications, products, Services, or models that have potential competitive possibilities with the Services without authorization. (6) Copying, transferring, renting, lending, selling, or providing sub-licensing or re-licensing of the Services in whole or in part without authorization. (7) Other activities that endanger network security, the operational security or create risks of the Services. **3.3 Maliciously harm or circumvent our safeguard for information and risk prevention of the Services in the following ways:** (1) Using variants, random characters, homophones, or other means to evade safety detection and input or generate illegal statements. (2) Launching malicious attacks, disseminating computer virus or inducing the Services. (3) Deleting, altering, or concealing the identification marks of artificially generated content that we have labeled, including both explicit and implicit. (4) Other behaviors that maliciously harm our management for information security and risk prevention of the Services. **3.4 Use the Services to:** (1) Use subliminal, manipulative, or deceptive techniques that distort a person's behavior so that they are unable to make informed decisions in a way that is likely to cause harm. (2) Exploit any vulnerabilities related to age, disability, or socio-economic circumstances to distort and harm. (3) Evaluate or classify individuals based on their social behavior or personal traits leading to detrimental or unfavorable treatment. (4) Assess or predict the risk of an individual committing a criminal offense based solely on their personal traits or on profiling. (5) Infer an individual's emotions in the workplace and educational settings, except when necessary for medical or safety reasons. (6) Conduct real-time remote biometric identification in public spaces for law enforcement purposes. (7) Create or expand facial recognition databases without consent. (8) Send us any personal information of children under 14 or the applicable age of digital consent or allow minors to use the Services without consent from their parent or guardian. (9) Use any method to extract data from the Services other than as permitted through the APIs; or buy, sell, or transfer API keys from, to or with a third party. If you use the Services to process personal data, you are required to provide legally adequate privacy notices and obtain all necessary consents for the processing of personal data by the Services, and process personal data in compliance with all applicable laws. You commit to not utilizing the Services for the creation, reception, maintenance, transmission, or any other processing of information that encompasses or qualifies as "Protected Health Information" as delineated by the HIPAA Privacy Rule. You acknowledge and agree that in the event of your violation of this agreement or applicable laws and regulations, we have the right to independently determine, including making a comprehensive judgment based on your personal information, behavioral data, and other interactive relationships. Without prior notice, we may take measures against you, including but not limited to warning, asking for rectification, restricting account functions, suspending use, freezing and confiscating the recharged amount, closing accounts, prohibiting re-registration, and deleting content. We have the right to announce the results of such actions and may decide whether to restore the use based on the actual circumstances. ## 4.Content You and End Users may submit prompts, texts, audios or other content or materials (**"input"**) to the Services and the Services generate corresponding content as a response (**"output"**). Both input and output are collectively referred to as **"content".** You represent and warrant that you own or have the necessary license, authorization or clearance to submit your input to the Services.You are solely responsible for content and we do not claim ownership of it. Due to technical limitations, we cannot guarantee that the content of other customers will be entirely different from yours, and it is possible that there may be similarities in the output. We may use Content to provide, maintain, develop, support, and improve the Services, comply with applicable law, enforce our terms and policies, and keep the Services safe and secure. Customer who requires restrictions on the use of Customer Content for training or improving Moonshot AI models may contact Moonshot AI to discuss available enterprise arrangements or separate written agreements. Unless otherwise expressly agreed in writing, Customer Content may be used for the foregoing purposes. Given the uncertainty and probabilistic nature of machine learning, we cannot guarantee the accuracy of output. Additionally, our output may not be able to exhaustively satisfy your desire. Therefore, please note the following when using the Services: (1) Do not regard the output as the sole source of fact or absolutely factual information. You need to assess the accuracy of the content.(2) The output should not be a substitute for professional advice in fields such as medical, legal, financial, and educational. (3) You must not use any output relating to a person for any purpose that could have a legal or material impact on that person, such as making credit, educational, employment, housing, insurance, legal, medical, or other important decisions about them. (4) The output does not represent the views of Moonshot AI. If the output mentions any third-party products or Services, it does not imply that the third party endorses the content or is affiliated with Moonshot AI. While we perform regular backups of Content, it does not guarantee that there will be no loss or corruption of data. Corrupt or invalid backup points may arise due to various factors, including but not limited to content that is already corrupted before the backup process begins or that changes during the backup operation.We will provide support and endeavor to troubleshoot any known or discovered issues that may impact the backups of content. However, you acknowledge that we have no liability related to the integrity of the content or the failure to successfully restore the content to a usable state.You agree to maintain a complete and accurate copy of any content in a location that is independent of the service. ## 5.Fees and Payment **5.1 Payment Obligation.** Customer agrees to pay all fees applicable to its use of the Services in accordance with the pricing set forth on the pricing page or in an applicable Order Form. Unless otherwise specified, all fees are due and payable using the payment method selected by Customer. **Payment Authorization.** By providing a payment method (such as a credit or debit card), Customer authorizes Moonshot AI and its payment processors to charge such payment method on a recurring and/or usage-based basis for all fees incurred in connection with the Services. **Billing and Auto-Charge.** Customer agrees that Moonshot AI may automatically charge the designated payment method for applicable fees, including usage-based charges, subscription fees (if any), and applicable taxes, as such fees become due. Customer acknowledges that usage-based charges may vary depending on actual use of the Services. **Failure of Payment.** If a charge is declined or fails for any reason, Moonshot AI may retry the charge, suspend access to the Services, or require Customer to provide an alternative payment method. **5.2** **Payment Method Management.** Customer is responsible for maintaining accurate and up-to-date payment information. Customer may update or remove its payment method at any time through its account settings, subject to any outstanding payment obligations. **Pricing Changes.** Moonshot AI may update pricing from time to time. Any updated pricing will apply to Customer’s use of the Services after the effective date of such change. For Customer subject to an Order Form, pricing shall remain as set forth in the applicable Order Form for its term, unless otherwise agreed in writing. **Promotions.** We may offer you promotions through means such as gifting recharge amounts, issuing vouchers, or providing trial Services. We reserve the right to decide whether to continue offering or to discontinue these Services. When using these Services, you should be aware of this characteristic to avoid any inconvenience caused by the discontinuation of such Services. The amounts associated with free Services are non-withdrawable, non-transferable, and non-invoiceable. **Taxes.** Fees are exclusive of applicable taxes. Customer is responsible for any taxes associated with its use of the Services, excluding taxes based on Moonshot AI’s net income. **Refunds.** Except as otherwise provided in this Agreement or required by applicable law, fees are non-refundable. Refunds may be issued where required by law or where Moonshot AI determines, in its reasonable discretion, that a refund is appropriate, including where Services are terminated due to legal requirements or material failure of the Services attributable to Moonshot AI. Customer may submit refund requests in accordance with the procedures made available by Moonshot AI from time to time. Moonshot AI will review such requests in good faith. **Billing Disputes.** Customer must notify Moonshot AI in writing of any good faith dispute of fees within thirty (30) days after the applicable charge, or such charge will be deemed accepted. **5.3 Account Responsibility and Authorized Persons.** The Enterprise User (the “User”) shall be responsible for all acts conducted under its User Account (including but not limited to signing agreements online by clicking “Agree,” checking boxes, or submitting information, as well as browsing, uploading, or inputting any content). All operations performed through the User Account are deemed duly authorized by the User and represent the User’s true intent, and are legally binding upon the User. Any person who logs into the Platform using the User Account and password and performs operations is deemed to have been fully authorized by the User (the “Authorized Person”). The Authorized Person shall have the authority to handle, on behalf of the User, matters including but not limited to the following: (1) sign, acknowledge and accept agreements, orders, confirmation letters and other legal documents relating to the Service; (2) purchase, activate, renew, upgrade or modify the Service and pay the corresponding fees on behalf of the User; (3) apply for, acknowledge and receive invoices; (4) confirm the delivery, acceptance, deployment, activation and usage of the Products or Services; and (5) handle other matters relating to the User’s use of the Service. The Authorized Person may use WeChat Pay, Alipay, bank cards, credit cards or other lawful payment accounts held in the name of the Authorized Person or a third party to perform payment obligations on behalf of the User. All such payments are deemed made on behalf of the User and do not alter the User’s status as the contracting party and payment obligor, nor do they create any independent contractual relationship between the payer of such payment (the “Payer”) and us. The User acknowledges and agrees that all amounts paid by the Authorized Person or other Payer on behalf of the User shall be applied to satisfy the User’s payment obligations under the relevant agreements and orders. Except where a refund is required by law due to our breach, or where this Agreement cannot be performed due to Force Majeure, the User shall not, on the grounds that the payment account is not the User’s own account, that the Payer is the User’s employee, affiliate, legal representative, shareholder, actual controller or other third party, or that there exists any internal dispute, authorization dispute, labor dispute, termination of entrustment or agency relationship or any other reason between the Payer and the User, claim the invalidity of the payment, demand rescission of the transaction, refuse to perform contractual obligations, or require us to refund the paid amounts to the Payer or any other third party. If we have reasonable grounds to believe that a payment is abnormal, or suspected unauthorized use of a payment account/card, impersonation, fraud, money laundering, commercial bribery or other risks of illegality or regulatory non-compliance, or if the User, Authorized Person or Payer fails to provide the authorization relationship, payment basis or other reasonable supporting materials as required by us, we may suspend the provision of Products or Services, refuse to accept the payment, suspend delivery, require a change of payment method, or terminate the relevant transaction. The User shall bear all liabilities arising therefrom. The User shall ensure that the payment activities of the Authorized Person and the Payer comply with applicable laws and regulations and do not involve money laundering, terrorist financing, commercial bribery, fraud, tax evasion or other illegal or criminal activities. If the payment activities of the User, Authorized Person or Payer violate laws or regulations, or if any third-party claim, administrative investigation, judicial proceeding or other dispute arises due to the payment authorization, source of funds, payment method or any other reason, causing us to suffer any losses, liabilities or expenses (including but not limited to attorneys’ fees, litigation costs, arbitration costs, investigation costs and compensation to third parties), the User shall indemnify and hold us harmless against all such losses, liabilities and expenses. **Special Notice.** The User acknowledges that it is solely responsible for managing and regulating the authority of its employees, management personnel, agents and other persons authorized to use the Account. Any payment, purchase, renewal, upgrade, service activation or other act by the Authorized Person is deemed an act of the User, and the User shall bear all legal consequences arising therefrom. The User may not assert against us any internal management reasons, including but not limited to incomplete internal approval procedures, absence of authorization documents, employee resignation, or acting beyond the scope of authority. ## 6.Confidential Information Each party ("Discloser") may disclose confidential or proprietary information to the other party ("Recipient") in connection with the Services, including technical, business, financial, and other non-public information that is identified as confidential or that reasonably should be understood to be confidential under the circumstances ("Confidential Information"). Customer Content shall be deemed Customer's Confidential Information. **Protection of Confidential Information.** The Recipient shall use the Confidential Information of the Discloser only as necessary to exercise its rights and perform its obligations under this Agreement. Recipient shall not disclose Confidential Information to any third party except to its employees, contractors, affiliates, and professional advisors who have a need to know such information and who are bound by confidentiality obligations at least as protective as those set forth herein. Recipient shall protect Confidential Information using at least reasonable care. **Exclusions.** Confidential Information does not include information that: (1) is or becomes publicly available through no fault of Recipient; (2) was lawfully known to Recipient without restriction before disclosure; (3) is lawfully received from a third party without restriction; or (4) is independently developed without the use of Confidential Information. **Required Disclosure.** Recipient may disclose Confidential Information to the extent required by applicable law, regulation, or court order, provided that, where legally permitted, Recipient gives reasonable prior notice to the Discloser and reasonably cooperates with efforts to limit the disclosure. ## 7.Intellectual Property Except for the relevant right holders who are entitled to rights in accordance with the law, Moonshot AI owns all rights (including but not limited to copyright, trademark rights, patent rights, and other intellectual property rights and additional rights) within the scope permitted by applicable laws and regulations with respect to the Services (including but not limited to software, technology, programs, code, models, user interfaces, web pages, text, charts, layout designs, trademarks, and electronic documents). ## 8.Warranties and Disclaimer We warranty that, during the term of this Agreement, the Services will substantially comply with the documentation we provide to you or make publicly available, when used in accordance with this Agreement. Apart from the warranties explicitly stated above, the Services are provided on an "as-is" basis. We, along with our affiliates and licensors, hereby disavow any and all other warranties, whether expressed or implied, including but not limited to implied warranties of merchantability, fitness for a particular purpose, title, non-infringement, and quiet enjoyment, as well as any warranties that may arise from the course of dealing or trade usage. Irrespective of any contrary statements, we do not make any representations or warranties regarding: (1) the uninterrupted, error-free, or secure use of the Services, (2) the correction of any defects, (3) the accuracy of content. ## 9.Limitation of liability To the fullest extent permitted by law, neither party, its affiliates, licensors, or suppliers shall be liable under or in connection with this Agreement for any consequential, indirect, special, incidental, exemplary, or punitive damages, including loss of profits, business, revenue, goodwill, anticipated savings, or data, even if advised of the possibility of such damages. To the fullest extent permitted by law, the maximum aggregate liability of Moonshot AI and its affiliates arising out of or relating to the Services or this Agreement shall not exceed the total amount paid by Customer to Moonshot AI for the Services during the twelve (12) months preceding the event giving rise to the claim. The foregoing exclusions and limitations shall not apply to: (1) a party's gross negligence or willful misconduct; (2) Customer's payment obligations; (3) either party's indemnification obligations; (4) breach of confidentiality obligations; or (5) violations of applicable data protection or privacy laws. ## 10.Indemnity You shall be responsible for and shall defend, indemnify, and hold harmless Moonshot AI and its affiliates, personnel, and agents from and against any third-party claims, damages, liabilities, costs, and expenses (including reasonable attorneys' fees) arising out of or relating to: (1) your or your End Users' use of the Services in violation of this Agreement or applicable law; (2) your Inputs or Customer Content, including where you do not have sufficient rights, consents, or permissions to provide such Inputs or Customer Content, or where such Inputs or Customer Content infringe or violate the rights of any third party; or (3) your combination, modification, or use of the Services or Outputs in a manner not authorized under this Agreement. Notwithstanding the foregoing, you shall not be responsible to the extent that any such claim or loss is finally determined to have been caused by Moonshot AI's material breach of this Agreement, gross negligence, or willful misconduct. ## 11.Termination or Discontinuation of the Services **Cancellation.** Customer may terminate the Services at any time by deleting your account.You may apply to delete your account by contacting us via our designated email address. Before you proceed with account cancellation, please be mindful of the balance in your account. Once your account is canceled, your account information, data, APIs, and any remaining balance will be permanently deleted. Even if you register again using the same entity, these items will not be restored. Therefore, please make your decision with caution. It will take us some time to review your cancellation request. When we receive your request to cancel your account, we may suspend the use of your account. If you wish to withdraw your cancellation request during the review period, please contact us promptly.During the cancellation process, if there is still a balance in your account that has not been used, we will send you a reminder to clear the balance. If you agree to proceed with the cancellation, we will continue with the cancellation process. If you do not explicitly agree, your account will not be cancelled for the time being. If you violate applicable laws and regulations, breaches these terms, or infringes upon the legitimate rights and interests of the public, us, or any third party, we may terminate the Services depending on the circumstances. For operational reasons, we reserve the right to decide on service settings and scope, and to suspend or discontinue the Services based on specific situations. After your service is terminated, we will delete your content and data in accordance with the requirements of applicable laws and regulations. ## 12.Governing Law and Dispute Resolution YOU AND US AGREE TO THE FOLLOWING PROVISIONS: The establishment, effectiveness, interpretation, revision, supplementation, termination, enforcement, and dispute resolution of these terms shall all be governed by the laws of Singapore. In the event of any dispute arising out of these terms, both you and us shall endeavor to resolve the dispute through amicable consultations. Engaging in this informal dispute resolution process is a requirement that must be completed before filing any legal action. If no resolution is reached through consultation in 60 days. Any dispute arising out of or in connection with these terms, including any question regarding existence, validity or termination of these terms, shall be referred to and finally resolved by arbitration administered by the Singapore International Arbitration Centre ("SIAC") in accordance with the Arbitration Rules of the Singapore International Arbitration Centre ("SIAC Rules") for the time being in force, which rules are deemed to be incorporated by reference in this clause. The seat of the arbitration shall be Singapore. The language of arbitration shall be English. YOU ACKNOWLEDGE AND AGREE THAT ANY LEGAL PROCEEDING OR ACTION RELATED TO THESE TERMS MUST BE INITIATED WITHIN ONE YEAR FROM THE DATE OF THE EVENT OR FACTS THAT GIVE RISE TO THE DISPUTE. FAILURE TO DO SO WILL RESULT IN A PERMANENT WAIVER OF YOUR RIGHT TO PURSUE ANY CLAIM OR CAUSE OF ACTION, REGARDLESS OF ITS NATURE OR TYPE, BASED ON SUCH EVENTS OR FACTS. ## 13.Miscellaneous **Changes to the Agreement.** We may update the Agreement from time to time as required. When we do, we will publish an updated version and effective date on the page, or by providing any other notice required by applicable law. We recommend that you review the updated agreements or terms each time you access the Services. **Assignment.** Customer may not assign, transfer, or delegate any of its rights or obligations under this Agreement without Moonshot AI's prior written consent, except to an affiliate or in connection with a merger, acquisition, or sale of substantially all of its assets, provided that the assignee agrees in writing to be bound by this Agreement. Moonshot AI may assign or transfer this Agreement to an affiliate or in connection with a merger, acquisition, or sale of all or substantially all of its assets, upon notice to Customer.Any attempted assignment in violation of this section shall be null and void. **Export and Sanctions.** Customer may not use, export, re-export, transfer, or provide access to the Services in violation of any applicable export control or sanctions laws or regulations, including those of the United States, Singapore, the European Union, and other applicable jurisdictions. Without limiting the foregoing, Customer may not use or provide access to the Services (1) in or to any country or region subject to comprehensive sanctions, or where such use would require an export license that has not been obtained, or (2) to any person or entity listed on any government restricted or denied party list. ## 14.Contact Us For feedback, appeal or complaint(especially copyright complaint) ,please contact us at [api-service@moonshot.ai](mailto:api-service@moonshot.ai). # Kimi OpenPlatform Privacy Policy Source: https://platform.kimi.ai/docs/agreement/userprivacy Learn how Kimi Open Platform collects, uses, stores, shares, and protects personal information and how users can exercise their data rights. **Last Update : April 30, 2025** Welcome to Kimi OpenPlatform! Our services are provided and controlled by MOONSHOT AI PTE. LTD. ("we"or"Moonshot AI") in Singapore through web pages.("the services") . The Kimi OpenPlatform Privacy Policy("this policy") is to clearly explain how we collect, use, disclose, and protect your personal information, as well as the rights you have and to help you make appropriate choices. Your privacy matters to us. Please do take the time to get to know and familiarize yourself with the policy and practices. By continuing to use the services, you agree with the policy. If you do not wish your personal information collected, used, or disclosed as described below, or if you are under the age of 18, you should stop accessing the services. We would like to specifically remind you that when you access third-party products and services through the services, the handling of your personal information and privacy will be managed by the third party in accordance with its own policies. We cannot be responsible for the processing activities of third parties. In such cases, we recommend that you carefully read the relevant policies of the third parties to understand your rights and obligations. ## 1. Personal Information We Collect ### Personal Information You Provide: **Account Information:** The information you provide when registering for and using an account, such as username, telephone number, profile picture, email address and account credentials. This information is used to create and manage your account, ensuring that you can access and use our services smoothly. **User content:** The prompts, audios, images, videos, files and other content that you input and generate while using our products and services. This information helps us optimize our models and understand your needs and preferences, so that we can provide you with more accurate services and support. **Communication Information:** If you communicate with us, such as via email or feedback pages or our pages on social media sites, we may collect personal information like your name, contact information, and the contents of the messages you send. **Surveys, Research, and Promotions:** We collect information you provide if you choose to participate in a survey, research, promotion, contest, marketing campaign, or event conducted or sponsored by us. ### We Receive from Your Use of the Services: We may receive the following information automatically when you visit, use or interact with our services: **Log Data:** We collect information that your browser or device automatically sends when you use the services. **Device and Usage Information:** We collect technical information related to your device and usage of our services. This includes device models, operating system versions, unique device identifiers, MAC addresses, user IDs, conversation IDs, IP addresses, browser types, telecommunications operators, language preferences, clipboard data, and access dates and times. The specific information collected may vary based on your device type and settings. **Cookies & Similar Technologies:** We may use cookies and similar technologies on our website to understand your usage patterns and provide more personalized services. If you use the services without creating an account, we may store some of the information described in this policy with cookies. We will obtain your consent to our use of cookies where required by applicable law. ### Personal Information we collect from third parties We may receive the information described in this policy from other sources, such as: **Technical Information:** If you choose to register or log in using a third-party service such as Google, we may collect information from that service, such as access tokens and public information like your profile picture and nickname, with your consent. This is for account binding with our services, enabling direct login. We may also request additional information if any required registration details are missing. **Security Information:** We may receive information from our trusted partners, such as security partners, to protect against fraud, abuse, and other security threats to our services. **Public Information:** We may obtain publicly available information via Internet sources in order to develop the models that power our services. ## 2. How we use Personal Information We use your information to provide, maintain, improve and develop the services, including for the following purposes: **To Provide the Services:** We use your information to create and manage your account, provide a seamless and convenient login experience, and enable you to interact with the services, such as engaging in conversations and accessing various features. **To Improve and Develop Our Services:** We analyze how you use the services to identify areas for improvement, develop new features, and enhance the overall user experience. This includes training and refining our underlying technology, such as machine learning models and algorithms. **To Communicate with You:** We may contact you to inform you of changes to the services, provide customer support, dealing with complaints or send you other service-related messages. **To Ensure the Safety and Security of Our Services:** We use your information to detect and prevent fraudulent activities, abuse, and other security threats. This helps us maintain the integrity and stability of our platform. **To Comply with Legal Obligations and Rules:** We use your information to comply with legal obligations, our policy and to protect the rights, privacy, safety, or property of our users, Kimi, or third parties. **To Promote Our Services:** We may use your information to promote our services or third-party services through marketing communications, contests, or promotions, with your consent where required. ## 3. How we share Personal Information We may share your information in the following circumstances: **Service Providers:** We engage service providers that help us provide, support, and develop the services and understand how they are used. We share your information with these service providers as necessary to enable them to provide their services. These providers will access, process, or store personal information only in the course of performing their duties to us. These providers include hosting services, customer service vendors, cloud services, content delivery services, support and safety monitoring services, email communication software, web analytics services, payment and transaction processors, and other information technology providers. **Affiliates:** We may share your information with our affiliates, ensuring that they handle your information in a manner consistent with this Privacy Policy. **Corporate Transactions:** In the event of a corporate transaction, such as a merger, sale of assets or shares, reorganization, financing, change of control, or acquisition of all or a portion of our business, your information may be shared with third parties. If such a transaction involves the transfer of your personal information, we will require the new holder of your personal information to continue to be bound by this policy. Otherwise, we will require the company, organization, or individual to obtain your consent again. **Legal Obligations and Rights:** We may share your information with government authorities or other third parties if we have good faith belief that it is necessary to: (i) comply with applicable law, legal process or government requests, as consistent with international standards; (ii) protect and defend the rights, property, and safety of us, users and public as well as the safety, security of our products, employees, users; (iii) if we determine, in our sole discretion, that there is a violation of our terms, policies, or the law; (iv) detect or prevent fraud or other illegal activity; (v) protect against legal liability. ## 4. Your Rights and Choices You have the rights to access, rectify or update, transfer, delete your personal information as well as restrict how we process, withdraw your consent, lodge complaints before the competent authorities and potentially others. You can exercise some of these rights directly through settings or through your account if you have registered. Additionally, your device may offer settings that allow you to manage the information we collect. For instance, you can decide whether to permit us to access your mobile advertising identifier through the settings on your Apple or Android device. If you prefer not to receive marketing or advertising emails, you can simply use the "unsubscribe" link or mechanism provided in those emails.If you are unable to exercise your rights through the methods mentioned above, please submit your request through [**api-service@moonshot.ai**](mailto:api-service@moonshot.ai). If you choose to delete your account, you will not be able to reactivate your account or retrieve any of the content or information in connection with your account. **A note about accuracy:** The services generate responses by analyzing a user's request and predicting the words that are most likely to follow. However, the most probable words may not always be factually accurate. As a result, you should not assume that the output from our models is factually correct. If you find that our output contains inaccurate information about you and you wish to request a correction or removal of that information, please contact us. We will review your request in accordance with applicable laws and the technical capabilities of our models. ## 5. Information Security We are fully aware that the security and confidentiality of your personal information are of utmost importance to you. We will endeavor to avoid collecting irrelevant user information and will take reasonable technical and organizational measures to protect your personal information from unauthorized access, public disclosure, use, modification, damage, or loss, including but not limited to: **Data Encryption:** We will use industry-leading encryption algorithms to encrypt users' personal information, ensuring that even if the data is obtained by unauthorized third parties, they will not be able to decipher the actual information of the users. **System Security Protection:** We regularly conduct security checks and vulnerability patches on servers and systems to prevent hacker attacks and virus intrusions. **Institutional Safeguards:** We have established a dedicated data protection department and appointed a person in charge to be responsible for the protection of users' personal information. We regularly provide training and education on personal information protection to our employees to ensure that all staff understand and value the protection of users' personal information. We strictly limit employees' access to user data, allowing only specific employees to access relevant data when necessary for business purposes. We also monitor and assess employees who access users' personal information, and take prompt action if any improper behavior is detected. **Data Backup:** We regularly back up users' personal information data and store the backup data in different physical locations to reduce the risk of data loss due to accidents or disasters. **Data Breach Alert:** We have set up a real-time monitoring system to monitor abnormal data access and leakage situations. Once signs of data leakage are detected, we will immediately take measures to stop the leakage, promptly identify the cause, and develop corresponding improvement measures. Please note that due to the limitations of technical and management means, although we have taken the above measures and made every effort to improve the security level, we cannot guarantee absolute security. Also, we cannot guarantee the effectiveness of privacy settings or security measures on third parties related to the services. **Therefore, you should take special care in deciding what information you send to us.** We strongly recommend that you take more proactive security measures together with us, such as refusing to lend accounts, disclose verification codes, or upload sensitive information, which are high-risk operations. We have developed a cybersecurity incident emergency response plan. In the unfortunate event of a personal information leak or other security incidents, we will immediately activate the emergency response plan to take measures to prevent the expansion of harm. We will promptly inform you of the basic situation of the security incident, the potential impact, and the remedial measures we have taken through means such as telephone or push notifications. When it is difficult to notify each individual whose personal information is affected, we will issue a public announcement in a reasonable and effective manner. At the same time, we will also report the relevant situation to the regulatory authorities, cooperate with the investigation, and hold the responsible parties accountable for their legal responsibilities. ## 6. Retention We store your information as long as necessary to provide the services, fulfill the purposes outlined in this policy and other legitimate business purposes (such as service improvement, resolving disputes, safety or security), comply with legal obligations. Retention periods vary based on factors such as information type, sensitivity, and legal requirements. For example, account, input, and payment information are retained while your account is active. Violation-related information is retained until the violation is resolved. Also, in some cases, the length of time we retain data depends on your settings. Your personal information may be transferred to and stored on servers located outside your country of residence. We store the information we collect in secure servers located in Singapore. When cross-border transfers are necessary, we will implement appropriate safeguards, consistent with applicable data protection laws, to ensure your information is protected for the purposes outlined in this policy. ## 7. Children Our Services are not directed to, or intended for, children under 18. If you have reason to believe that a child under 14 has provided information to us through the services, please contact [api-service@moonshot.ai](mailto:api-service@moonshot.ai). We will investigate any notification and, if appropriate, delete the information in time. ## 8. Privacy Policy Updates We may update this policy from time to time as required. When we do, we will publish an updated version and effective date on this page, or by providing any other notice required by applicable law. We recommend that you review the updated policy each time you access the services to stay informed of our privacy practices. ## 9. Contact us Please send detailed comments, complaints, or reports concerning this policy or our personal information management practices to [api-service@moonshot.ai](mailto:api-service@moonshot.ai). ## 10. Supplemental Terms - Jurisdiction-Specific **European Economic Area ("EEA"), Switzerland, and UK** If you are using the services in the EEA, Switzerland or the UK (the "European Region"), the following additional terms apply: **Legal bases for processing** We use your personal data only as permitted by law. Our legal bases for processing your personal data described in this policy are described in the table below. | Purpose of processing | Personal information categories | Legal bases | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | To create and manage your account, provide a seamless and convenient login experience, and enable you to interact with our services, such as engaging in conversations and accessing various features. | - Account Information
- User Content
- Communication Information
- Device and Usage Information
- Log Data
- Cookies & Similar Technologies | - **Where necessary to perform a contract** with you when we provide and maintain our services.
- **Your consent** when we ask for it to process your personal data for a specific purpose to help you create accounts.
- **Where necessary for our legitimate interests:** If you're under 18, your ability to enter into a binding contract is restricted. In such cases, where we may not be able to process your information based on contractual necessity, we rely on legitimate interests. We use the information we collect to enable you to access and use the platform. | | To identify areas for improvement, develop new features, and enhance the overall user experience. This includes training and refining our underlying technology, such as machine learning models and algorithms. | - User Content
- Communication Information
- Device and Usage Information
- Log Data | - **Where necessary to pursue our legitimate interests in** identifying and resolving issues with the platform and enhancing its functionality; and to conduct research and safeguard personal data through aggregation or anonymization, in line with the principles of data minimization and privacy by design.
- **Your consent** when we ask for it to process your personal data for a specific purpose to provide services. | | To inform you of changes to our services, provide customer support, or send you other service-related messages. | - Account Information
- Communication Information
- Device and Usage Information | - **Where necessary to perform a contract** with you when we provide and maintain our services.
- **Your consent** when we ask for it to process your personal data for a specific purpose to communicate with you. | | To comply with our legal obligations, or as necessary to perform tasks in the public interest, or to protect the vital interests of our users and other people. This could include providing law enforcement agencies or emergency services with information in urgent situations to protect health or life. | - Account Information
- User Content
- Communication Information
- Device and Usage Information
- Log Data
- Cookies & Similar Technologies | - **Where necessary to comply with a legal obligation:** This applies in situations where we are required to take steps to ensure the safety of our users, or to comply with a valid legal request, such as an order from law enforcement agencies or courts. It also applies when we are fulfilling our obligations under applicable laws, or when we are protecting the rights, safety, or property of ourselves, our affiliates, our users, or third parties.
- **Where necessary for our legitimate interests** to protect our business, employees, and users from illegal activities, inappropriate conduct, or breaches of terms that could cause harm, and to disclose and share information with regulators or other government authorities. | | To detect and prevent fraudulent activities, abuse, and other security threats. This helps us maintain the integrity and stability of our platform. | - Account Information
- User Content
- Communication Information
- Device and Usage Information
- Log Data
- Cookies & Similar Technologies | - **Where necessary to comply with a legal obligation** when we use your personal data to comply with applicable law or when we protect our or our affiliates', users', or third parties' rights, safety, and property.
- **Where necessary for our legitimate interests** to ensure that the Platform is safe and secure. | **Data storage** It is important to note that our servers are situated in Singapore. Consequently, when you utilize our services, your personal data may be processed and stored on our servers located in Singapore. This could involve either a direct submission of your personal data to us or a transfer of your data conducted by us or a third-party. **Your rights** You have the following rights: * The right to request free of charge (i) confirmation of whether we process your personal data and (ii) access to a copy of the personal data retained; * The right to request proper rectification or erasure of your personal data or restriction of the processing of your personal data; * Where processing of your personal data is either based on your consent or necessary for the performance of a contract with you and processing is carried out by automated means, the right to receive the personal data concerning you in a structured, commonly used and machine-readable format or to have your personal data transmitted directly to another company, where technically feasible (data portability); * Where the processing of your personal data is based on your consent, the right to withdraw your consent at any time (withdrawal will not impact the lawfulness of data processing activities that have taken place before such withdrawal); * The right not to be subject to any automatic individual decisions, including profiling, which produces legal effects on you or similarly significantly affects you unless we have your consent, this is authorised by European Union or Member State law or this is necessary for the performance of a contract; * The right to object to processing if we are processing your personal data on the basis of our legitimate interest unless we can demonstrate compelling legitimate grounds which may override your right. If you object to such processing, we ask you to state the grounds of your objection in order for us to examine the processing of your personal data and to balance our legitimate interest in processing and your objection to this processing; * The right to request the restriction of the processing of your information where (i) you are challenging the accuracy of the information, (ii) the information has been unlawfully processed, but you are opposing the deletion of that information, (iii) you need the information to be retained for the pursuit or defense of a legal claim, or (iv) you have objected to the processing and you are awaiting the outcome of that objection request; * The right to object to processing your personal data for direct marketing purposes; and * The right to lodge complaints before your local data protection authority. Before we can respond to a request to exercise one or more of the rights listed above, you may be required to verify your identity or your account details. Please contact us if you would like to exercise any of your rights. # Check Balance Source: https://platform.kimi.ai/docs/api/balance GET /v1/users/me/balance REST API to check your available, voucher, and cash balances on Kimi OpenPlatform. Check your available, voucher, and cash balances on the Kimi Open Platform. When the available balance is less than or equal to 0, you cannot call the inference API. ```python python expandable theme={null} import os import requests api_key = os.environ.get("MOONSHOT_API_KEY") url = "https://api.moonshot.ai/v1/users/me/balance" response = requests.get( url, headers={"Authorization": f"Bearer {api_key}"}, ) print(response.json()) ``` ```bash curl expandable theme={null} curl https://api.moonshot.ai/v1/users/me/balance \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```javascript node.js expandable theme={null} const apiKey = process.env.MOONSHOT_API_KEY; async function main() { const response = await fetch("https://api.moonshot.ai/v1/users/me/balance", { method: "GET", headers: { Authorization: `Bearer ${apiKey}`, }, }); const data = await response.json(); console.log(data); } main(); ``` | Field | Type | Description | | ------------------------ | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `code` | integer | Response code. 0 indicates success. | | `data` | object | Balance data object | | `data.available_balance` | number | The available balance (unit: USD), including cash balance and voucher balance. When it is less than or equal to 0, the user cannot call the inference API | | `data.voucher_balance` | number | The voucher balance (unit: USD), which cannot be negative | | `data.cash_balance` | number | The cash balance (unit: USD), which can be negative, indicating that the user owes money. When it is negative, `available_balance` is equal to the value of `voucher_balance` | | `scode` | string | Status code | | `status` | boolean | Request status | **Response Example** ```json theme={null} { "code": 0, "data": { "available_balance": 49.58894, "voucher_balance": 46.58893, "cash_balance": 3.00001 }, "scode": "0x0", "status": true } ``` When `available_balance` is less than or equal to 0, API requests will return an `exceeded_current_quota_error` error. Please recharge your account or check your voucher validity. API Keys from `platform.kimi.ai` and `platform.kimi.com` are completely independent. Using a Key from one platform on the other will result in a 401 error. Ensure your endpoint matches the platform where the Key was created. # Cancel Batch Source: https://platform.kimi.ai/docs/api/batch-cancel POST /v1/batches/{batch_id}/cancel Cancel an in-progress batch task. The status will change to cancelling and then to cancelled. Only tasks in validating, in_progress, or finalizing status can be cancelled. Cancel an in-progress batch task. After cancellation, the task status will change to `cancelling` and then to `cancelled`. Only tasks in `validating`, `in_progress`, or `finalizing` status can be cancelled. ```python Python theme={null} import os from openai import OpenAI from openai.types import Batch client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) batch: Batch = client.batches.cancel("your_batch_id") print(f"Status: {batch.status}") # cancelling ``` ```bash cURL theme={null} curl -X POST ${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/batches/your_batch_id/cancel \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { const batch = await client.batches.cancel("your_batch_id"); console.log(`Status: ${batch.status}`); // cancelling } main(); ``` On success, the endpoint returns a `BatchObject` containing the following fields: | Field | Type | Description | | ------------------- | --------------- | -------------------------------------------------------------------------------------------------------------------------- | | `id` | string | Unique identifier for the batch task | | `object` | string | Object type, always `batch` | | `endpoint` | string | Request endpoint | | `input_file_id` | string | Input file ID | | `completion_window` | string | Processing time window | | `status` | string | Current status. After a successful cancellation request this is typically `cancelling`, and eventually becomes `cancelled` | | `output_file_id` | string \| null | Output file ID for successful results | | `error_file_id` | string \| null | Error file ID for failed results | | `created_at` | integer | Creation timestamp (Unix) | | `in_progress_at` | integer \| null | Execution start timestamp (Unix) | | `expires_at` | integer \| null | Expiration timestamp (Unix) | | `finalizing_at` | integer \| null | Result preparation start timestamp (Unix) | | `completed_at` | integer \| null | Completion timestamp (Unix) | | `failed_at` | integer \| null | Validation failure timestamp (Unix) | | `cancelling_at` | integer \| null | Cancellation request timestamp (Unix) | | `cancelled_at` | integer \| null | Cancellation completion timestamp (Unix) | | `request_counts` | object | Request counts, including `completed`, `failed`, and `total` | | `metadata` | object \| null | Custom metadata | Only tasks in `validating`, `in_progress`, or `finalizing` status can be cancelled. If the task is already `completed`, `failed`, `expired`, or `cancelled`, calling this endpoint will return a 400 error. **Common Errors** * **400 Bad Request**: The task status does not allow cancellation, or the request parameters are invalid. Please verify the task status before attempting cancellation. * **401 Unauthorized**: Invalid or missing API key. Check that `Authorization: Bearer ` is correct. * **404 Not Found**: The specified `batch_id` does not exist. Verify the ID is correct and the task belongs to your organization. * **500 Server Error**: Internal server error. Please retry later; if the issue persists, contact support with the `request_id`. For more details, see [Error Codes](/docs/api/errors). For complete usage examples and status transitions, see the [Batch API Guide](/docs/guide/use-batch-api). # Create Batch Source: https://platform.kimi.ai/docs/api/batch-create POST /v1/batches Create a batch task. You need to first upload a JSONL file with purpose="batch" via the Files API, then use the returned file_id to create the task. **Limits:** | Limit | Description | | ------------------ | ------------------------------------------------------------ | | File format | Must have `.jsonl` extension | | File size | Must be non-empty, max 100MB | | Organization quota | Up to 1000 batch-purpose files per organization | | Model consistency | All requests in a batch must use the same model | | `custom_id` | Must be unique within the file | | Model access | The specified model must exist and the user must have access | For complete usage examples, see the [Batch API Guide](/docs/guide/use-batch-api). # List Batches Source: https://platform.kimi.ai/docs/api/batch-list GET /v1/batches List batch tasks for your organization. List all batch tasks for your organization, with pagination support. Commonly used to review task status or manage batch operations. ```python Python theme={null} import os from openai import OpenAI from openai.pagination import SyncCursorPage from openai.types import Batch client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) batches: SyncCursorPage[Batch] = client.batches.list(limit=10) for batch in batches.data: print(f"{batch.id} - {batch.status} ({batch.request_counts.completed}/{batch.request_counts.total})") ``` ```bash cURL theme={null} curl "${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/batches?limit=10" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { const batches = await client.batches.list({ limit: 10 }); for (const batch of batches.data) { console.log(`${batch.id} - ${batch.status} (${batch.request_counts.completed}/${batch.request_counts.total})`); } } main(); ``` The list batches endpoint returns a paginated list object containing the following fields: | Field | Type | Description | | ---------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `object` | string | Object type, always `"list"` | | `data` | array\[object] | List of batch tasks, each element is a `BatchObject` with the same fields as [Retrieve Batch](/docs/api/batch-retrieve) | | `has_more` | boolean | Whether there are more results. When `true`, pass the last `batch.id` from this page as the `after` parameter to fetch the next page | **Response Example** ```json theme={null} { "object": "list", "data": [ { "id": "batch_xxx", "object": "batch", "endpoint": "/v1/chat/completions", "input_file_id": "file_xxx", "completion_window": "24h", "status": "completed", "output_file_id": "file_yyy", "error_file_id": null, "created_at": 1711475054, "in_progress_at": 1711475055, "expires_at": 1711561454, "finalizing_at": 1711475100, "completed_at": 1711475110, "failed_at": null, "cancelling_at": null, "cancelled_at": null, "request_counts": { "completed": 100, "failed": 0, "total": 100 }, "metadata": null } ], "has_more": false } ``` For complete usage examples and status transitions, see the [Batch API Guide](/docs/guide/use-batch-api). When `has_more` is `true`, you must use the `after` parameter for pagination. Pass the last `batch.id` from the current page as the `after` value to retrieve the next page. Omitting `after` will always return the first page. # Retrieve Batch Source: https://platform.kimi.ai/docs/api/batch-retrieve GET /v1/batches/{batch_id} Retrieve the status and details of a specific batch task. Retrieve the current status, progress statistics, and detailed metadata of a specific batch task. Typically used to poll whether a task has completed after creation. ```python Python theme={null} import os from openai import OpenAI from openai.types import Batch client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) batch: Batch = client.batches.retrieve("your_batch_id") print(f"Status: {batch.status}") print(f"Progress: {batch.request_counts.completed}/{batch.request_counts.total}") ``` ```bash cURL theme={null} curl ${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/batches/your_batch_id \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { const batch = await client.batches.retrieve("your_batch_id"); console.log(`Status: ${batch.status}`); console.log(`Progress: ${batch.request_counts.completed}/${batch.request_counts.total}`); } main(); ``` | Field | Type | Description | | ------------------- | --------------- | ---------------------------------------------------------------------------------------------------------------------- | | `id` | string | Unique identifier for the batch task | | `object` | string | Object type, always `batch` | | `endpoint` | string | Request endpoint | | `input_file_id` | string | Input file ID | | `completion_window` | string | Processing time window | | `status` | string | Current status: `validating`, `failed`, `in_progress`, `finalizing`, `completed`, `expired`, `cancelling`, `cancelled` | | `output_file_id` | string \| null | Output file ID for successful results | | `error_file_id` | string \| null | Error file ID for failed results | | `created_at` | integer | Creation timestamp (Unix) | | `in_progress_at` | integer \| null | Execution start timestamp (Unix) | | `expires_at` | integer \| null | Expiration timestamp (Unix) | | `finalizing_at` | integer \| null | Result preparation start timestamp (Unix) | | `completed_at` | integer \| null | Completion timestamp (Unix) | | `failed_at` | integer \| null | Validation failure timestamp (Unix) | | `cancelling_at` | integer \| null | Cancellation request timestamp (Unix) | | `cancelled_at` | integer \| null | Cancellation completion timestamp (Unix) | | `request_counts` | object | Request count statistics, containing `completed`, `failed`, and `total` | | `metadata` | object \| null | Custom metadata | For complete usage examples and polling scripts, see the [Batch API Guide](/docs/guide/use-batch-api). When `status` is `completed`, `output_file_id` contains the results file ID. When `status` is `failed`, check `error_file_id` for error details. If the specified `batch_id` does not exist, the API returns a `404` error (`resource_not_found_error`). # Create Chat Completion Source: https://platform.kimi.ai/docs/api/chat POST /v1/chat/completions Creates a completion for the chat message. Supports standard chat, Partial Mode, and Tool Use (Function Calling). Create a chat completion request. The model generates a response based on the provided message list. The `content` field supports the following two forms: **Plain text string** ```json theme={null} { "content": "Hello" } ``` **Array of objects** (for multimodal input) Each element in the array is distinguished by the `type` field: ```json theme={null} { "content": [ { "type": "text", "text": "Describe this image" }, { "type": "image_url", "image_url": { "url": "data:image/png;base64,..." } }, { "type": "video_url", "video_url": { "url": "data:video/mp4;base64,..." } } ] } ``` `image_url` and `video_url` also support passing a string directly, equivalent to the `url` field in object form: ```json theme={null} { "type": "image_url", "image_url": "data:image/png;base64,..." } ``` #### Parameter Description Each element in the array has the following fields: | Parameter | Required | Description | Type | | ----------- | ------------------------------ | --------------------------------------------------------------------------------------- | ------------------------------------------ | | `type` | required | Content type | `"text"` \| `"image_url"` \| `"video_url"` | | `text` | required when `type=text` | Text content | string | | `image_url` | required when `type=image_url` | For transmitting images. Supports object form `{"url": "..."}` or a URL string directly | object \| string | | `video_url` | required when `type=video_url` | For transmitting videos. Supports object form `{"url": "..."}` or a URL string directly | object \| string | When `image_url` is passed as an object, its fields are: | Parameter | Required | Description | Type | | --------- | -------- | ------------------------------------------------------ | ------ | | `url` | required | Image content specified via base64 encoding or file id | string | When `video_url` is passed as an object, its fields are: | Parameter | Required | Description | Type | | --------- | -------- | ----------------------------------------------------------------------------------------------- | ------ | | `url` | required | Video content specified via base64 encoding or file id, for example `data:video/mp4;base64,...` | string | Both the object form (`url` field) and the string shorthand support the following formats: * Base64 encoding: `data:image/png;base64,...` or `data:video/mp4;base64,...` * File reference: `ms://` See [Use the Kimi Vision Model](/docs/guide/use-kimi-vision-model). #### Usage Example ```python python expandable theme={null} import os import base64 from openai import OpenAI from openai.types.chat import ChatCompletion client: OpenAI = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1", ) # Encode the image to base64 with open("your_image_path", "rb") as f: img_base: str = base64.b64encode(f.read()).decode("utf-8") response: ChatCompletion = client.chat.completions.create( model="kimi-k2.6", messages=[ { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": f"data:image/jpeg;base64,{img_base}", }, }, { "type": "text", "text": "Describe this image", }, ], } ], ) print(response.choices[0].message.content) ``` ```bash curl expandable theme={null} curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k2.6", "messages": [ { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "data:image/jpeg;base64,/9j/4AAQ..." } }, { "type": "text", "text": "Describe this image" } ] } ] }' ``` ```javascript node.js expandable theme={null} const fs = require("fs"); const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }); async function main() { // Encode the image to base64 const imgBase = fs.readFileSync("your_image_path").toString("base64"); const response = await client.chat.completions.create({ model: "kimi-k2.6", messages: [ { role: "user", content: [ { type: "image_url", image_url: { url: `data:image/jpeg;base64,${imgBase}`, }, }, { type: "text", text: "Describe this image", }, ], }, ], }); console.log(response.choices[0].message.content); } main(); ``` ### Non-streaming Response ```json theme={null} { "id": "cmpl-04ea926191a14749b7f2c7a48a68abc6", "object": "chat.completion", "created": 1698999496, "model": "kimi-k2.6", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Hello, Li Lei! 1+1 equals 2. If you have any other questions, feel free to ask!" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 19, "completion_tokens": 21, "total_tokens": 40, "cached_tokens": 10 } } ``` ### Streaming Response ```text theme={null} data: {"id":"cmpl-xxx","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k2.6","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]} data: {"id":"cmpl-xxx","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k2.6","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]} ... data: {"id":"cmpl-xxx","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k2.6","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":19,"completion_tokens":13,"total_tokens":32,"cached_tokens":12}} data: [DONE] ``` The model name in the response example will be returned based on the model parameter in the request. When using the `kimi-k2.6` model, the `"model"` field in the response will show `"kimi-k2.6"`. The Kimi API is stateless and does not retain conversation history. To implement multi-turn dialogue, append the previous assistant reply (and any tool results, if applicable) back into the `messages` array before sending the next request. ```python theme={null} messages = [ {"role": "system", "content": "You are Kimi."}, {"role": "user", "content": "Hello, my name is Li Lei."} ] completion = client.chat.completions.create(model="kimi-k2.6", messages=messages) reply = completion.choices[0].message # Append the assistant reply back into messages for the next turn messages.append({"role": "assistant", "content": reply.content}) messages.append({"role": "user", "content": "What is 1+1?"}) ``` When the conversation history grows too long, retain only the most recent messages or compress earlier turns to avoid exceeding the model's context limit. See [Engage in Multi-turn Conversations](/docs/guide/engage-in-multi-turn-conversations-using-kimi-api) for details. Use the `response_format` parameter to constrain the model output format: * `{"type": "text"}` (default): plain text output * `{"type": "json_object"}`: forces a valid JSON Object output * `{"type": "json_schema", "json_schema": {...}}`: outputs structured data according to the given JSON Schema (Structured Output) When using `json_object`, **you must explicitly describe the expected JSON fields and types in the system prompt or user prompt**, otherwise the model may produce unexpected results. ```json theme={null} { "model": "kimi-k2.6", "messages": [ {"role": "system", "content": "Please output JSON containing title, author, and summary fields."}, {"role": "user", "content": "Summarize this article..."} ], "response_format": {"type": "json_object"} } ``` See [Use the JSON Mode Feature of Kimi API](/docs/guide/use-json-mode-feature-of-kimi-api) for details. Pass external tools defined as JSON Schema via the `tools` parameter. The model can decide to invoke them when appropriate. **Request example** ```json theme={null} { "model": "kimi-k2.6", "messages": [{"role": "user", "content": "What is the weather in Beijing today?"}], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Get the weather for a given city", "parameters": { "type": "object", "properties": { "city": {"type": "string", "description": "City name"} }, "required": ["city"] } } } ] } ``` **`tool_calls` in the response** When `finish_reason` is `"tool_calls"`, the model returns a `tool_calls` array containing `id`, `function.name`, and `function.arguments`: ```json theme={null} { "choices": [{ "message": { "role": "assistant", "content": "", "tool_calls": [{ "id": "call_xxx", "type": "function", "function": { "name": "get_weather", "arguments": "{\"city\":\"Beijing\"}" } }] }, "finish_reason": "tool_calls" }] } ``` **Submitting tool execution results** After executing the tool locally, append the result back into `messages` using `role="tool"`. The `tool_call_id` must match the `id` from the request: ```json theme={null} {"role": "tool", "tool_call_id": "call_xxx", "content": "Sunny, 25°C"} ``` See [Use the Kimi API to Complete Tool Calls](/docs/guide/use-kimi-api-to-complete-tool-calls) for details. `kimi-k3` always reasons and uses the top-level `reasoning_effort` field (`"low"`, `"high"`, or `"max"`; default `"max"`). `kimi-k2.6` and `kimi-k2.7-code` support thinking mode: the model first outputs its reasoning process (`reasoning_content`) before producing the final answer. **K2.x request parameters** | Field | Type | Description | | --------------- | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `thinking.type` | `"enabled"` \| `"disabled"` | Thinking switch (`kimi-k2.7-code` is always `enabled` and cannot be disabled) | | `thinking.keep` | `null` \| `"all"` | Preserved Thinking: whether to retain historical `reasoning_content` in context. `kimi-k2.6` defaults to `null` (not kept) and accepts `"all"`; `kimi-k2.7-code` is fixed at `"all"` (kept), other values return an error | **Response fields** In non-streaming responses, `choices[0].message` contains: | Field | Description | | ------------------- | --------------------------------------------------------------- | | `content` | Final answer | | `reasoning_content` | Reasoning process (returned only when thinking mode is enabled) | ```json theme={null} { "choices": [{ "message": { "role": "assistant", "content": "1+1 equals 2.", "reasoning_content": "The user asked a basic math question, simply add the numbers." } }] } ``` When using a thinking model in multi-turn conversations, always preserve the `reasoning_content` of each historical assistant message in `messages`, otherwise the model may lose reasoning context. See [Use the Kimi K2 Thinking Mode](/docs/guide/use-thinking-models) for details. Set `stream: true` to enable streaming output. The model returns content incrementally in Server-Sent Events (SSE) format. Recommended for scenarios requiring real-time feedback, such as chat, code generation, and long text output. ```python theme={null} completion = client.chat.completions.create( model="kimi-k2.6", messages=[{"role": "user", "content": "Explain what recursion is."}], stream=True ) for chunk in completion: if chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="") ``` **SSE Response Format** Each line starts with `data:`, followed by a JSON object. When `finish_reason` is `null`, content accumulates in `delta.content`; when `finish_reason` is not `null`, the output is complete: ```text theme={null} data: {"id":"cmpl-xxx","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k2.6","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]} data: {"id":"cmpl-xxx","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k2.6","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]} data: {"id":"cmpl-xxx","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k2.6","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":19,"completion_tokens":13,"total_tokens":32,"cached_tokens":12}} data: [DONE] ``` **`stream_options`** Use `stream_options: {"include_usage": true}` to receive an additional `usage` field in the last chunk (before `data: [DONE]`), showing the token consumption of the request: ```python theme={null} stream=True, stream_options={"include_usage": True} ``` See [Utilize the Streaming Output Feature of Kimi API](/docs/guide/utilize-the-streaming-output-feature-of-kimi-api) for details. Partial Mode (Prefill) allows you to prefill an output prefix in the last assistant message of `messages`, guiding the model to continue generation in the format or direction you expect. **How to Enable** Append an `role="assistant"` message at the end of the `messages` array, and set `partial: true`: ````python theme={null} completion = client.chat.completions.create( model="kimi-k2.6", messages=[ {"role": "user", "content": "Implement quicksort in Python."}, {"role": "assistant", "content": "```python\n", "partial": True} ] ) ```` The model will continue generating code from \`\`\`\`python\n\` instead of outputting explanatory text first. **Common Use Cases** * Force the model to start with a specific format (e.g., JSON's `{`, a code block's \`\`\`\`python\`) * Maintain role name prefixes in role-play scenarios (combined with the `name` field) * When `finish_reason="length"`, use the same prefix to continue truncated content Do not mix Partial Mode with `response_format={"type": "json_object"}`, as this may lead to unexpected model responses. To guide JSON output, use [Structured Output](/docs/guide/response_format) directly, or set `partial: true` and prefill `{` separately. See [Use the Partial Mode Feature of Kimi API](/docs/guide/use-partial-mode-feature-of-kimi-api) for details. # Common Error Codes Source: https://platform.kimi.ai/docs/api/errors Look up Kimi API HTTP status codes, error types, common causes, and recommended troubleshooting steps. When a request fails, the API returns a JSON response that contains error information: ```json theme={null} { "error": { "type": "content_filter", "message": "The request was rejected because it was considered high risk" } } ``` ## Error List ## 400 — Bad Request | error type | Typical message | Cause and fix | | ----------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | | `content_filter` | The request was rejected because it was considered high risk | The input or model output triggered content safety review. Modify the prompt and avoid sensitive or high-risk content. | | `invalid_request_error` | Request format error, missing required parameter, or invalid parameter type | Check the request body against the API documentation. | | `invalid_request_error` | Input token length too long | The input tokens exceed the model's maximum context limit. Shorten the input or use a model with a larger context window. | | `invalid_request_error` | prompt tokens + max\_tokens exceeds the model specification | Reduce `max_tokens` or switch to another model. | | `invalid_request_error` | Invalid purpose: only 'file-extract' accepted | The `purpose` field for file upload is incorrect. Currently, only `file-extract` is supported. | | `invalid_request_error` | File size is too large, max file size is 100MB, please confirm and re-upload the file | The uploaded file exceeds the 100MB limit. Compress or split the file and upload it again. | | `invalid_request_error` | File size is zero, please confirm and re-upload the file | The uploaded file size is 0. Check whether the file is corrupted or empty. | | `invalid_request_error` | Too many uploaded files | The total number of uploaded files exceeds the limit. Delete earlier files that are no longer used, then try again. | ## 401 — Authentication Error | error type | Typical message | Cause and fix | | ------------------------------ | -------------------------- | ------------------------------------------------------------------------- | | `invalid_authentication_error` | Invalid Authentication | The API key is invalid or malformed. Check `Authorization: Bearer `. | | `incorrect_api_key_error` | Incorrect API key provided | The API key was not provided, or the key is incorrect. | **Platform key isolation**: Keys issued on `platform.kimi.ai` are independent from keys issued on other regional Kimi platforms. Mixing keys across platforms returns 401. Make sure the endpoint matches the platform where the key was created. ## 403 — Permission Error | error type | Typical message | Cause and fix | | ------------------------- | -------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- | | `permission_denied_error` | The API you are accessing is not open | This API is not currently available to your account. | | `permission_denied_error` | You are not allowed to get other user info | You are not allowed to access other users' information. Check the permission scope for the API. | | `permission_denied_error` | Your IP is not allowed to access this organization | The calling IP is not in the organization's allowlist. This is common on the international platform. Contact an administrator to add the IP. | ## 404 — Resource Not Found | error type | Typical message | Cause and fix | | -------------------------- | ----------------------------------------------------------------------------- | ----------------------------------------------------------------- | | `resource_not_found_error` | Model not found, or this account does not have permission to access the model | Check the spelling of the `model` parameter and the account tier. | ## 429 — Rate Limit / Insufficient Quota | error type | Typical message | Cause and fix | | ------------------------------ | ---------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `engine_overloaded_error` | The engine is currently overloaded, please try again later | The service node is under high load (for example, peak-hour capacity pressure). Wait as indicated by `Retry-After`, reduce concurrency, and retry with exponential backoff. This is caused by server-side capacity; topping up or upgrading your tier does not resolve it. | | `exceeded_current_quota_error` | Account balance is insufficient or the account has been disabled | Check your balance and billing status. | | `exceeded_current_quota_error` | Token quota is insufficient | Top up your account and try again. | | `rate_limit_reached_error` | Organization-level concurrency limit reached | Reduce concurrency or retry after the time indicated in the response. | | `rate_limit_reached_error` | Organization-level RPM limit reached | Retry after waiting for the time indicated in the response. RPM means requests per minute. | | `rate_limit_reached_error` | Organization-level TPM limit reached | Reduce request frequency or upgrade your tier. TPM means tokens per minute. | | `rate_limit_reached_error` | Organization-level TPD limit reached | The limit will reset the next day, or you can upgrade your plan. TPD means tokens per day. | ## 499 / 500 / 503 / 504 — Connection And Server Errors | HTTP | error type | Cause and fix | | ---- | ------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 499 | `client_closed_request` | The client disconnected before the server returned a response. This is common when a streaming response is cut off by an intermediate proxy or when the user cancels the request. Check KeepAlive and timeout settings. | | 500 | `server_error` / `unexpected_output` | Internal server error. Try again later. If the issue persists, contact support with the `request_id`. | | 503 | `server_unavailable` | The service is temporarily unavailable. Try again later. This is usually related to node scaling or maintenance. | | 504 | `504 Gateway Time-out` | No response from the server for 900 seconds; the gateway returns an HTML timeout page. Common with long non-streaming requests. Use streaming output (`stream: true`). | ## Troubleshooting Tips * **401 response**: First confirm that you are using an API key from the correct platform. * **429 response**: First identify the cause by `error.type`: back off and retry for node overload, reduce concurrency or upgrade your account tier for organization-level rate limits, and top up for insufficient balance. See [Top-up and Rate Limits](/docs/pricing/limits). * **500 response**: Try again later. If the issue persists, contact the support team at [api-service@moonshot.ai](mailto:api-service@moonshot.ai). * **504 response**: The gateway timed out because the server produced no response for 900 seconds. Use streaming output (`stream: true`). # Estimate Tokens Source: https://platform.kimi.ai/docs/api/estimate POST /v1/tokenizers/estimate-token-count Estimates the number of tokens that would be used for a given set of messages and model. The input structure is almost identical to that of chat completion. The input structure for `estimate-token-count` is almost identical to that of `chat completion`. ```bash theme={null} curl 'https://api.moonshot.ai/v1/tokenizers/estimate-token-count' \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k3", "messages": [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You excel in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You refuse to answer any questions involving terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages." }, { "role": "user", "content": "Hello, my name is Li Lei. What is 1+1?" } ] }' ``` ```python theme={null} import os import base64 import json import requests api_key = os.environ.get("MOONSHOT_API_KEY") endpoint = "https://api.moonshot.ai/v1/tokenizers/estimate-token-count" image_path = "image.png" with open(image_path, "rb") as f: image_data = f.read() # Encode the image to base64 format for the image_url image_url = f"data:image/{os.path.splitext(image_path)[1].lstrip('.')};base64,{base64.b64encode(image_data).decode('utf-8')}" payload = { "model": "kimi-k3", "messages": [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You excel in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You refuse to answer any questions involving terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": image_url, }, }, { "type": "text", "text": "Please describe the content of the image.", }, ], } ] } response = requests.post( endpoint, headers={ "Authorization": f"Bearer {api_key}", "Content-Type": "application/json" }, data=json.dumps(payload) ) print(response.json()) ``` If there is no `error` field, you can take `data.total_tokens` as the calculation result. # Files Source: https://platform.kimi.ai/docs/api/files Explore Kimi API file management for uploading, listing, retrieving, reading, and deleting files used for content extraction or visual understanding. The Kimi API provides file management capabilities for content extraction, image understanding, and video analysis. Upload files for content extraction or vision understanding List all files uploaded by the current user Retrieve metadata for a specific file Delete files that are no longer needed Retrieve extracted file content # Get File Content Source: https://platform.kimi.ai/docs/api/files-content GET /v1/files/{file_id}/content Retrieves extracted text content for files uploaded with purpose `file-extract`. ```python theme={null} # Note: retrieve_content is marked with a warning in the latest version. Use the line below instead. # If you are using an older version, you can use retrieve_content. file_content = client.files.content(file_id=file_object.id).text ``` ```bash theme={null} curl https://api.moonshot.ai/v1/files/{file_id}/content \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` # Delete File Source: https://platform.kimi.ai/docs/api/files-delete DELETE /v1/files/{file_id} Deletes a previously uploaded file. Deletes a previously uploaded file. Once deleted, the file no longer counts against your storage quota and cannot be used in chat completions or Batch requests. ```python showLineNumbers expandable theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1", ) # Delete a specific file delete_result = client.files.delete(file_id="file-xxx") print(delete_result) ``` ```bash showLineNumbers theme={null} curl -X DELETE https://api.moonshot.ai/v1/files/file-xxx \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```js showLineNumbers expandable theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }); async function main() { const deleteResult = await client.files.delete({ file_id: "file-xxx", }); console.log(deleteResult); } main(); ``` The delete file endpoint returns a JSON object with the following fields: | Field | Type | Description | | --------- | ------- | ----------------------------------------- | | `id` | string | Deleted file identifier | | `object` | string | Always `"file"` | | `deleted` | boolean | Whether the file was deleted successfully | **Response Example** ```json theme={null} { "id": "file-xxx", "object": "file", "deleted": true } ``` Deletion is permanent. Each user can upload a maximum of 1,000 files. When the total file count reaches the limit, you must delete unused files before uploading new ones. If the file does not exist or has already been deleted, the endpoint returns a `404` error. Please verify that the `file_id` is correct. # List Files Source: https://platform.kimi.ai/docs/api/files-list GET /v1/files Lists all files uploaded by the current user. ```python theme={null} file_list = client.files.list() for file in file_list.data: print(file) # Inspect the metadata of each file ``` # Get File Information Source: https://platform.kimi.ai/docs/api/files-retrieve GET /v1/files/{file_id} Retrieves metadata for a specific uploaded file. ```python theme={null} client.files.retrieve(file_id=file_id) # FileObject( # id='clg681objj8g9m7n4je0', # bytes=761790, # created_at=1700815879, # filename='xlnet.pdf', # object='file', # purpose='file-extract', # status='ok', # status_details='' # ) ``` # Upload File Source: https://platform.kimi.ai/docs/api/files-upload POST /v1/files Uploads a file for extraction, image understanding, or video understanding. Each user can upload a maximum of 1,000 files. Each file must not exceed 100 MB, and the total size of all uploaded files must not exceed 10 GB. The file parsing service is currently free, but rate limiting may be applied during peak traffic periods. Supported formats include `.pdf`, `.txt`, `.csv`, `.doc`, `.docx`, `.xls`, `.xlsx`, `.ppt`, `.pptx`, `.md`, `.jpeg`, `.png`, `.bmp`, `.gif`, `.webp`, `.ico`, `.xbm`, `.dib`, `.pjp`, `.tif`, `.pjpeg`, `.avif`, `.dot`, `.apng`, `.epub`, `.tiff`, `.jfif`, `.html`, `.json`, `.mobi`, `.log`, `.go`, `.h`, `.c`, `.cpp`, `.cxx`, `.cc`, `.cs`, `.java`, `.js`, `.css`, `.jsp`, `.php`, `.py`, `.py3`, `.asp`, `.yaml`, `.yml`, `.ini`, `.conf`, `.ts`, `.tsx`, and more. When uploading a file, use `purpose="file-extract"` if you want the model to use the extracted file contents as context. ```python showLineNumbers expandable theme={null} from pathlib import Path import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url = "https://api.moonshot.ai/v1", ) # xlnet.pdf is an example file; we support pdf, doc, and image formats. file_object = client.files.create(file=Path("xlnet.pdf"), purpose="file-extract") # Note: retrieve_content is deprecated in the latest version. # If you are using the latest SDK, use files.content instead. file_content = client.files.content(file_id=file_object.id).text messages = [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are particularly skilled in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You will refuse to answer any questions involving terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages.", }, { "role": "system", "content": file_content, }, {"role": "user", "content": "Please give a brief introduction of what xlnet.pdf is about"}, ] completion = client.chat.completions.create( model="kimi-k2-turbo-preview", messages=messages, temperature=0.6, ) print(completion.choices[0].message) ``` ```bash showLineNumbers theme={null} # xlnet.pdf is a sample file curl https://api.moonshot.ai/v1/files \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -F purpose="file-extract" \ -F file="@xlnet.pdf" ``` ```js showLineNumbers expandable theme={null} const OpenAI = require("openai"); const fs = require("fs"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }); async function main() { let file_object = await client.files.create({ file: fs.createReadStream("xlnet.pdf"), purpose: "file-extract" }); // retrieve_content is deprecated in the latest version. let file_content = await (await client.files.content(file_object.id)).text(); let messages = [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are more proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You will refuse to answer any questions related to terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages.", }, { "role": "system", "content": file_content, }, {"role": "user", "content": "Please give a brief introduction of what xlnet.pdf is about"}, ]; const completion = await client.chat.completions.create({ model: "kimi-k2-turbo-preview", messages: messages, temperature: 0.6 }); console.log(completion.choices[0].message.content); } main(); ``` Replace `$MOONSHOT_API_KEY` with your own API key, or set it as an environment variable before making the call. If you want to upload multiple files and have a conversation with Kimi based on these files, you can use the following pattern: ```python expandable theme={null} from typing import * import os import json from pathlib import Path from openai import OpenAI client = OpenAI( base_url="https://api.moonshot.ai/v1", api_key=os.environ["MOONSHOT_DEMO_API_KEY"], ) def upload_files(files: List[str]) -> List[Dict[str, Any]]: """ upload_files uploads all provided files (paths) via the file upload API '/v1/files', retrieves the uploaded file content, and generates file messages. Each file becomes an independent message with role set to system. The Kimi model will correctly recognize the file content in these system messages. :param files: A list of file paths to upload. Paths can be absolute or relative, passed as strings. :return: A list of messages containing file content. Add these messages to the Context, i.e., the messages parameter when calling the `/v1/chat/completions` API. """ messages = [] for file in files: file_object = client.files.create(file=Path(file), purpose="file-extract") file_content = client.files.content(file_id=file_object.id).text messages.append({ "role": "system", "content": file_content, }) return messages def main(): file_messages = upload_files(files=["upload_files.py"]) messages = [ *file_messages, { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are more proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You will refuse to answer any questions related to terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages.", }, { "role": "user", "content": "Summarize the content of these files.", }, ] print(json.dumps(messages, indent=2, ensure_ascii=False)) completion = client.chat.completions.create( model="kimi-k2-turbo-preview", messages=messages, ) print(completion.choices[0].message.content) if __name__ == '__main__': main() ``` When uploading image or video assets for native model understanding, use `purpose="image"` or `purpose="video"`. Please refer to [Using Vision Models](/docs/guide/use-kimi-vision-model) for end-to-end examples. # Join the Kimi Developer Community Source: https://platform.kimi.ai/docs/api/join-the-community Connect with Kimi developers on Discord and the developer forum. Connect with other Kimi developers, ask questions, share feedback, and showcase what you are building with Kimi. Chat with the community, get help, and keep up with the latest Kimi updates. Start discussions, explore community knowledge, and share ideas with other developers. # List Models Source: https://platform.kimi.ai/docs/api/list-models GET /v1/models List all currently available models. List all currently available models, including model ID, context length, and capability flags. We recommend querying this endpoint before creating a chat completion to verify the target model is available and supports the required capabilities. ```python python expandable theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1", ) models = client.models.list() print(models.data) ``` ```bash curl expandable theme={null} curl https://api.moonshot.ai/v1/models \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```javascript node.js expandable theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }); async function main() { const models = await client.models.list(); console.log(models.data); } main(); ``` The response is an object containing the following fields: | Field | Type | Description | | -------- | -------------- | ---------------------------------------------- | | `object` | string | Object type, always `"list"` | | `data` | array\[object] | List of models, each element is a model object | Each model object in the `data` array has the following fields: | Field | Type | Description | | -------------------- | ------- | ------------------------------------------------------ | | `id` | string | Model ID, e.g. `kimi-k3` | | `object` | string | Object type, always `"model"` | | `created` | integer | Unix timestamp when the model was created | | `owned_by` | string | Model owner identifier, e.g. `"moonshot"` | | `context_length` | integer | Maximum context length supported by the model (tokens) | | `supports_image_in` | boolean | Whether the model supports image input | | `supports_video_in` | boolean | Whether the model supports video input | | `supports_reasoning` | boolean | Whether the model supports deep thinking | The model list and capability flags may change as the platform is updated. If you receive a `resource_not_found_error` (404) when creating a chat completion, the model does not exist or your account does not have access to it. Check the `model` parameter spelling and confirm your account tier. This endpoint requires a valid API Key for authentication. If you receive a 401 error, verify that the `Authorization` header is `Bearer ` and that the API Key matches the platform endpoint (Keys from `platform.kimi.ai` and `platform.kimi.com` are not interchangeable). # Model Parameter Reference Source: https://platform.kimi.ai/docs/api/models-overview Compare default values, supported ranges, and constraints for Chat Completions API parameters across Kimi model families. Different model families have different defaults and constraints for Chat Completions API parameters. For the full model list, see the [Model List](/docs/models). ## Parameter Comparison When `temperature` is close to 0, `n` can only be 1. Otherwise, the API returns `invalid_request_error`. ## Model Parameter Differences When switching models, you need to look beyond the `model` field — models differ in which request parameters they support and what defaults they use: | Parameter | `kimi-k3` | `kimi-k2.7-code` | `kimi-k2.6` | `kimi-k2.5` | | ---------------------------------------- | ---------------------------------------------- | ------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | ----------------------------------------------------- | | Context window | 1M tokens | 256K tokens | 256K tokens | 256K tokens | | `thinking` | — | May be omitted; if set explicitly, only `{"type":"enabled","keep":"all"}` is accepted | `{"type":"enabled"}` (default), `{"type":"disabled"}`, `{"type":"enabled","keep":"all"}` | `{"type":"enabled"}` (default), `{"type":"disabled"}` | | `reasoning_effort` | `"low"` / `"high"` / `"max"` (default `"max"`) | Not supported | Not supported | Not supported | | `tool_choice` | `auto` / `none` / `required` | `required` not supported | `required` not supported | — | | `temperature` | Fixed at `1.0` | Fixed at `1.0` | `1.0` thinking / `0.6` non-thinking | `1.0` thinking / `0.6` non-thinking | | `top_p` | Fixed at `0.95` | Fixed at `0.95` | Fixed at `0.95` | — | | `n` | Fixed at `1` | Fixed at `1` | Fixed at `1` | — | | `presence_penalty` / `frequency_penalty` | Fixed at `0` | Fixed at `0` | Fixed at `0` | — | "Fixed" means the parameter cannot be modified: passing any other value returns an error, so do not pass it explicitly. ### `thinking` `thinking` is a K2.x-only request parameter: * `kimi-k2.6`: supports `{"type": "enabled"}` (default), `{"type": "disabled"}`, and `{"type": "enabled", "keep": "all"}`. * `kimi-k2.7-code`: thinking is on by default and only `{"type": "enabled", "keep": "all"}` is accepted; any other configuration returns an error. When switching from `kimi-k2.6`, you must pass back the historical `reasoning_content` in `messages` as required by Preserved Thinking. See [Thinking Mode](/docs/guide/use-thinking-models). ### `reasoning_effort` K3 always reasons with Preserved Thinking enabled. Configure its reasoning effort with the top-level `reasoning_effort` request field, which supports `"low"`, `"high"`, and `"max"` (default `"max"`). See [Reasoning Effort](/docs/guide/use-reasoning-effort). Switching levels invalidates prefix-cache hits. Decide on the `effort` level before the conversation starts and avoid switching it mid-session. ### `tool_choice` `kimi-k3` supports `auto` / `none` / `required`. `kimi-k2.6` and `kimi-k2.7-code` do not support `required` and return an error if it is passed. See [Tool Choice](/docs/guide/use-tool-choice). ### `temperature` * `kimi-k2.6` / `kimi-k2.5`: fixed at `1.0` in thinking mode and `0.6` in non-thinking mode; other values return an error. * `kimi-k2.7-code`: fixed at `1.0`; other values return an error. * `kimi-k3`: fixed at `1.0`; other values return an error. Do not pass `temperature` explicitly when calling these models. `kimi-k2.7-code-highspeed` is the same model as `kimi-k2.7-code` with identical parameter constraints; only the output speed differs. ### FAQ **Switching from `kimi-k2.6` to `kimi-k3` — do I need to change my code?** Replace `model` with `kimi-k3` and remove the K2.x `thinking` configuration. To set the reasoning effort explicitly, use top-level `reasoning_effort`. In multi-turn conversations and tool calls, pass the complete assistant message returned by the API back to `messages` as-is, including any `reasoning_content`. **Switching from `kimi-k2.7-code` to `kimi-k3` — do I need to change my code?** Replace `model` and continue passing complete assistant messages back as-is. To set the reasoning effort explicitly, use top-level `reasoning_effort`. **My code uses OpenAI's `reasoning_effort` — do I need to change it for `kimi-k3`?** No. K3 supports top-level `reasoning_effort` with `"low"`, `"high"`, and `"max"` as accepted values and `"max"` as the default. **Can I use `tool_choice: "required"` on `kimi-k2.6` or `kimi-k2.7-code`?** No. These models do not support `required` and return an error if it is passed; only `kimi-k3` supports it. ## Kimi K2.7 Code series — thinking Parameter The `kimi-k2.7-code` series includes `kimi-k2.7-code` and its high-speed variant `kimi-k2.7-code-highspeed`; the two are the same model with identical parameter constraints (including the table above and the `thinking` behavior) and differ only in output speed (referred to collectively as `kimi-k2.7-code` below). `kimi-k2.7-code` is code-focused, and all parameter constraints except `thinking` are identical to `kimi-k2.6`. Unlike `kimi-k2.6`, its **thinking is always on and cannot be disabled** (passing `{"type": "disabled"}` errors), and **Preserved Thinking is always on** (`thinking.keep` is treated as `"all"` whether omitted or set to `"all"`; any other invalid value errors). So you do not need to pass the `thinking` parameter — just switch the `model`, and the model always emits `reasoning_content`. For details, see [Using Thinking Mode](/docs/guide/use-thinking-models). ## Kimi K2.6 — thinking Parameter Kimi K2.6 supports the `thinking` parameter to control whether deep thinking is enabled. Accepts `{"type": "enabled"}` or `{"type": "disabled"}`. Since the OpenAI SDK doesn't have a native `thinking` parameter, use `extra_body`: ```python Python theme={null} completion = client.chat.completions.create( model="kimi-k2.6", messages=[ {"role": "user", "content": "Hello"} ], extra_body={ "thinking": {"type": "disabled"} }, max_tokens=1024*32, ) ``` ```bash cURL theme={null} curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k2.6", "messages": [ {"role": "user", "content": "Hello"} ], "thinking": {"type": "disabled"} }' ``` # API Overview Source: https://platform.kimi.ai/docs/api/overview Review Kimi API base URLs, authentication, request conventions, compatibility, and links to the main API endpoints. ## Service Address ``` https://api.moonshot.ai ``` Kimi Open Platform provides OpenAI-compatible HTTP APIs. You can use the OpenAI SDK directly. When using SDKs, set `base_url` to `https://api.moonshot.ai/v1`. When calling HTTP endpoints directly, use the full path such as `https://api.moonshot.ai/v1/chat/completions`. ## OpenAI Compatibility Our API is compatible with the OpenAI Chat Completions API in request/response format. This means: * You can use the official OpenAI SDKs (Python / Node.js) directly * Most OpenAI-compatible third-party tools and frameworks (LangChain, Dify, Coze, etc.) are supported * Simply point `base_url` to `https://api.moonshot.ai/v1` to switch Some parameters are Kimi-specific extensions: the `thinking` parameter needs to be passed via the SDK's `extra_body`; `partial` is a field on assistant messages within the messages array (`"partial": true`), not a top-level request parameter. See [Tool Use](/docs/api/tool-use) and [Partial Mode](/docs/api/partial) for details. ## Authentication All API requests require an API Key in the HTTP header: ``` Authorization: Bearer $MOONSHOT_API_KEY ``` API Keys can be created and managed in the [Kimi Open Platform Console](https://platform.kimi.ai/console/api-keys). Your API Key is sensitive. Do not expose it in client-side code, public repositories, or logs. Use environment variables to manage it. ## SDK Installation ```bash Python theme={null} pip install --upgrade 'openai>=1.0' ``` ```bash Node.js theme={null} npm install openai ``` Initialize the client: ```python Python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }); ``` Python version ≥ 3.7.1, Node.js version ≥ 18, OpenAI SDK version ≥ 1.0.0. ```bash theme={null} python -c 'import openai; print("version =", openai.__version__)' ``` ## Common Request Headers | Header | Value | Description | | --------------- | -------------------------- | -------------------- | | `Content-Type` | `application/json` | Request body format | | `Authorization` | `Bearer $MOONSHOT_API_KEY` | Authentication token | ## Error Handling When a request fails, the API returns a JSON error response containing `error.type` and `error.message` fields. Common HTTP status codes include 400 (bad request), 401 (authentication failure), 429 (rate limit), 500 (server error), etc. For the full list of error types, messages, and troubleshooting tips, see [Errors](/docs/api/errors). ## API Endpoints | Endpoint | Method | Description | | ------------------------------------- | ------ | -------------------------------------- | | `/v1/chat/completions` | POST | [Create Chat Completion](/docs/api/chat) | | `/v1/models` | GET | [List Models](/docs/api/list-models) | | `/v1/tokenizers/estimate-token-count` | POST | [Estimate Tokens](/docs/api/estimate) | | `/v1/users/me/balance` | GET | [Check Balance](/docs/api/balance) | | `/v1/files` | POST | [Upload File](/docs/api/files-upload) | | `/v1/files` | GET | [List Files](/docs/api/files-list) | | `/v1/files/{file_id}` | GET | [Get File Info](/docs/api/files-retrieve) | | `/v1/files/{file_id}` | DELETE | [Delete File](/docs/api/files-delete) | | `/v1/files/{file_id}/content` | GET | [Get File Content](/docs/api/files-content) | ## Next Steps Send your first API request Compare model capabilities and parameters Enable function calling Full endpoint parameter reference # Account and Billing Source: https://platform.kimi.ai/docs/guide/account-and-payments Find answers about Kimi Open Platform recharge, balances, credits, invoices, billing, and organization verification. * Individual users: Go to the user top-up page and complete an online payment. Online top-ups support WeChat Pay and Alipay QR code payments. After the payment succeeds, your account tier will be adjusted based on your accumulated top-up amount. * Business users: Please contact the Kimi sales team or follow the payment methods supported in your account. Supported options may include online payment or bank transfer, depending on your account region and billing setup. After the payment is received, your account tier will be adjusted based on your accumulated top-up amount. When using large models for code generation, the model may need multiple attempts to produce the expected code because of the randomness and complexity of generation. Programming tools may automatically perform multiple rounds of retries and calls, which can cause token usage to grow quickly. To better control costs and improve the usage experience, we recommend the following: * **Budget control** * **Set a daily spending limit**: Before use, go to [Kimi Open Platform project settings](https://platform.kimi.ai/console/projects/settings) and configure a project daily spending budget. Once the budget limit is reached, the system will automatically reject all API requests under that project. Note that due to billing latency, the limit may take about 10 minutes to take effect. For setup instructions, see [Organization Management Best Practices](/docs/guide/org-best-practice). * **Balance alerts**: We recommend enabling account balance alerts. When your account balance falls below the preset amount, the system will notify you so you can top up in time. * **Usage recommendations** * Start with a short context and a clear prompt for testing, then gradually add the complete business context. * **Continuous monitoring**: Keep monitoring your programming tool while it is running, handle abnormal situations promptly, and avoid unnecessary resource consumption caused by infinite loops or excessive retries. * **Model selection**: If cost is a concern, you can use the `kimi-k2.6` model. Kimi K3 offers a 1M-token context and uses flat pay-as-you-go pricing — there is no tiering by context length. Input (with separate rates for cache hits and misses) and output are billed at uniform per-token prices. See [Kimi K3 pricing](/docs/pricing/chat-k3). To ensure fair resource allocation and prevent abuse, rate limits are currently based on the account's accumulated top-up amount. For higher limits, please submit the [rate limit increase form](https://platform.kimi.ai/contact-sales). For more details, see the [Top-up and Rate Limits](https://platform.kimi.ai/docs/pricing/limits) page. * The Kimi sales team can provide additional resources and support for business customers using the Kimi API. Please fill out the form at [https://platform.kimi.ai/contact-sales](https://platform.kimi.ai/contact-sales) to contact sales. * Kimi Assistant now offers Kimi Business membership benefits. Visit [https://www.kimi.ai/membership/pricing](https://www.kimi.ai/membership/pricing) to subscribe online. * The platform supports issuing invoices based on either consumed amount or top-up amount. Please submit an invoice request online in [Invoice Management](https://platform.kimi.ai/console/invoice). * Supported invoice titles and requirements may vary by account type and billing setup. Please follow the options shown on the invoice request page. * Invoice type and tax rate: The issuing entity is Beijing Moonshot AI Technology Co., Ltd.; the service category is information technology services; the tax-inclusive rate is 6%. * Yes. After setting a password, you can log in with either your phone number and password or your account name and password. See [Account Password Settings](https://platform.kimi.ai/profile). * Yes. The target phone number must not have been registered on Kimi Open Platform or Kimi Assistant. * Account deletion is not currently supported. # Manage Email and Google Sign-In Source: https://platform.kimi.ai/docs/guide/account-security-and-sign-in Bind or update an email address, connect a Google account, and understand how sign-in methods affect your Kimi Platform account. Kimi Platform supports two account sign-in methods: * **Email sign-in:** Receive a verification code at your email address. After the address is added under **Email Binding**, it can also be used for password reset. * **Sign in with Google:** Authorize Kimi Platform through Google's OAuth sign-in flow. The connected identity appears under **Google Account Binding**. **Email Binding and Google Account Binding are separate credentials.** A Gmail address used for email verification is not automatically the same credential as a Google account with the same address. ## Open account security settings 1. Sign in to [Kimi Platform](https://platform.kimi.ai/). 2. Select **User Center** in the top navigation. You can also open [Account Settings](https://platform.kimi.ai/profile) directly. 3. Under **Security Information**, locate **Email Binding** and **Google Account Binding**. Select User Center in the Kimi Platform navigation From this section, you can manage the following settings: | Setting | Available actions | What it controls | | -------------------------- | ------------------------------------- | --------------------------------------------- | | **Email Binding** | **Bind Email** or **Modify** | Email verification sign-in and password reset | | **Google Account Binding** | **Bind Google Account** or **Unbind** | OAuth-based **Sign in with Google** | ## Bind or change an email address If no email address is bound: 1. Select **Bind Email**. 2. Enter the email address you want to use. 3. Request and enter the verification code sent to that address. 4. Complete verification. If an email address is already bound, select **Modify**. You will need to verify the current address and the new address. The new address must not already be bound to another Kimi Platform account. The following state shows an account with an email address bound and no connected Google account: An email address is bound and no Google account is connected ## Bind a Google account 1. Under **Google Account Binding**, select **Bind Google Account**. 2. Choose a Google account and complete Google's authorization flow. 3. Return to User Center and confirm that the Google account appears under **Google Account Binding**. After both methods are configured, you can sign in with the bound email address or use **Sign in with Google**. An email address and a Google account are both connected **Binding fails if the selected Google identity is already connected to another Kimi Platform account.** Choose a different Google account or remove that Google Account Binding from the other account first. ## How sign-in order affects account bindings The order in which you use email verification and Google sign-in can affect what appears in User Center. ### Email verification first, then Google When you first sign in with a verification code sent to an email address, the address appears under **Email Binding**. **Google Account Binding** remains empty. If you later use **Sign in with Google** with the Google account that has the same email address, Kimi Platform automatically adds that Google identity to the existing account, provided that: * The Kimi Platform account does not already have a different Google account bound. * The selected Google identity is not bound to another Kimi Platform account. ### Google first, then email verification When you first use **Sign in with Google**, the Google identity appears under **Google Account Binding**. **Email Binding** may remain empty. Signing in later with an email verification code sent to the same address does not automatically add that address under **Email Binding**. If User Center still shows only the Google account, select **Bind Email** to add an email credential explicitly. Bind an email address after signing in with Google ## Special case: the email address matches a different Google account Consider this configuration: * **Account 1** has **Email A** under Email Binding. * **Account 1** already has **Google Account B** under Google Account Binding. * **Google Account A** uses the same email address as **Email A**. If you try to use **Sign in with Google** with Google Account A, sign-in or binding may fail with a message indicating that the account is already linked to another third-party account. This happens because Account 1 already has Google Account B as its Google credential. The matching email address does not replace an existing Google Account Binding. To switch from Google Account B to Google Account A: 1. Sign in to Account 1 using Email A or Google Account B. 2. Open **User Center** → **Security Information**. 3. Confirm that Email A is listed under **Email Binding**. 4. Select **Unbind** next to Google Account B. 5. Select **Bind Google Account**, then authorize Google Account A. If Google Account A is already bound to another Kimi Platform account, it cannot be bound to Account 1 until it is removed from the other account. Kimi Platform does not merge accounts automatically based only on matching email addresses. ## Unbind a Google account Before unbinding your Google account, make sure an email address is configured under **Email Binding**. If Google is your only sign-in method, select **Bind Email** and complete email verification first. Then select **Unbind** under **Google Account Binding**. After unbinding, you can no longer use that Google identity to sign in to the account unless you bind it again. **Do not remove your only available sign-in method.** Keep a verified email address bound before disconnecting Google access. ## Frequently asked questions ### Is email verification with a Gmail address the same as Sign in with Google? No. Email verification proves access to an email inbox. **Sign in with Google** authorizes a Google identity through OAuth. They are managed as separate bindings even when the displayed email address is the same. ### Does changing Email Binding change Google Account Binding? No. Updating the bound email address does not automatically replace or remove the connected Google account. Manage each binding separately in User Center. ### Does a matching email address automatically merge two accounts? No. Kimi Platform does not merge accounts solely because an Email Binding and a Google identity display the same address. Existing account ownership and third-party bindings are checked before a Google identity can be connected. # AI-Readable Docs Source: https://platform.kimi.ai/docs/guide/ai-readable-docs Provide the Kimi API Platform documentation to AI coding assistants, enterprise bots, or RAG systems through llms.txt, llms-full.txt, the OpenAPI schema, and per-page Markdown. The Kimi API Platform documentation site provides several machine-readable endpoints. You can hand the entire documentation directly to AI coding assistants, enterprise bots, or RAG systems without crawling pages one by one. ## Site-wide endpoints | Endpoint | URL | Description | | :------------- | :----------------------------------------------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | llms.txt | [https://platform.kimi.ai/docs/llms.txt](https://platform.kimi.ai/docs/llms.txt) | An index of all documentation pages in a directory-style layout, about 9 KB. Start here to give a model the site structure, then let it fetch specific pages on demand | | llms-full.txt | [https://platform.kimi.ai/docs/llms-full.txt](https://platform.kimi.ai/docs/llms-full.txt) | The complete documentation in Markdown, about 700 KB. Suitable for loading as full context or as RAG source material | | OpenAPI Schema | [https://platform.kimi.ai/docs/openapi.json](https://platform.kimi.ai/docs/openapi.json) | The raw API specification (OpenAPI 3.1). Suitable for API integration and generating client code | ## Per-page Markdown Append `.md` to any documentation page URL to get that page in Markdown, for example [https://platform.kimi.ai/docs/overview.md](https://platform.kimi.ai/docs/overview.md). The menu at the top right of every page also offers Copy Page (copy this page as Markdown) and an option to open the page in an AI assistant, which works well when you only need specific pages. Copy page menu: Copy page, View as Markdown, and Open in Kimi ## Recommended usage * **AI coding assistants** (Claude Code, Cursor, and similar tools): add the llms.txt URL to your rules or context file. The model learns the site structure first, then fetches specific pages on demand; * **RAG indexing and enterprise bot knowledge bases**: crawl llms-full.txt to obtain the entire site in one pass, split it by page, and build your index; * **API integration and client code generation**: use openapi.json directly instead of extracting endpoint details from documentation pages. The Chinese documentation site offers the same endpoints. # Automatic Reconnection on Disconnect Source: https://platform.kimi.ai/docs/guide/auto-reconnect Implement automatic reconnection for Kimi API streaming requests and use Partial Mode to continue generation after an interruption. Due to concurrency limits, complex network environments, and other unforeseen circumstances, our connections may sometimes be interrupted. Typically, these intermittent disruptions don't last long. We want our services to remain stable even in such cases. Implementing a simple reconnection feature can be achieved with just a few lines of code. The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ```python theme={null} import os from openai import OpenAI import time client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url = "https://api.moonshot.ai/v1", ) def chat_once(msgs): response = client.chat.completions.create( model = "kimi-k3", messages = msgs ) return response.choices[0].message.content def chat(input: str, max_attempts: int = 100) -> str | None: messages = [ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are proficient in Chinese and English conversations. You aim to provide users with safe, helpful, and accurate answers. You will refuse to answer any questions related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated into other languages."}, ] # We construct the user's latest question as a message (role=user) and append it to the end of the messages list messages.append({ "role": "user", "content": input, }) st_time = time.time() for i in range(max_attempts): print(f"Attempts: {i+1}/{max_attempts}") try: response = chat_once(messages) ed_time = time.time() print("Query Successful!") print(f"Query Time: {ed_time-st_time}") return response except Exception as e: print(e) time.sleep(1) continue print("Query Failed.") return print(chat("Hello, please tell me a fairy tale.")) ``` The code above implements a simple reconnection feature, allowing up to 100 retries with a 1-second wait between each attempt. You can adjust these values and the conditions for retries based on your specific needs. # Best Practices for Benchmarking Source: https://platform.kimi.ai/docs/guide/benchmark-best-practice Run reproducible Kimi model benchmarks with recommended parameters, sample counts, streaming requests, and retry strategies. Benchmarking is an **engineering task** that needs stability and reproducibility. You'll be calling the model thousands of times; even tiny drifts in system setup or network latency can compromise result accuracy. Here's what we've learned to keep things reproducible and trustworthy. **Quick notes** * For any **unlisted** or **closed-source** benchmark: set`temperature = 1.0`, `stream = true`, `top_p = 0.95` * **Reasoning benchmarks**: `max_tokens = 128k`, and run at least **500–1000 samples** to get low variance (e.g. `AIME 2025`: 32 runs -> 30 × 32 = 960 questions) * **Coding benchmarks**: `max_tokens = 256k` * **Agentic task benchmarks:** * For multi-hop search: `max_tokens = 256k` + context management * Others: `max_tokens ≥ 16k–64k` ## K2.6 Models Benchmark Recommended Settings
Benchmark Category Benchmark Temperature Recommended max tokens Recommended runs Top-p Others (e.g. test log)
Multi-modal MMMU-Pro 1.0 max tokens = 96k 3 top\_p=0.95 thinking=
MMMU-Pro w/ python 1.0 per step tokens = 64k;
total max tokens = 256k
3 top\_p=0.95 Recommended max steps = 50
thinking=
CharXiv (RQ) 1.0 max tokens = 96k 3 top\_p=0.95 thinking=
CharXiv (RQ) w/ python 1.0 per step tokens = 64k;
total max tokens = 256k
3 top\_p=0.95 Recommended max steps = 50
thinking=
MathVision 1.0 max tokens = 96k 3 top\_p=0.95 thinking=
MathVision w/ python 1.0 per step tokens = 64k;
total max tokens = 256k
3 top\_p=0.95 Recommended max steps = 50
thinking=
V\* w/ python 1.0 per step tokens = 64k;
total max tokens = 256k
3 top\_p=0.95 Recommended max steps = 50
thinking=
Agent HLE-Full w/ tools 1.0 per step tokens = 48k;
total max tokens = 256k
1 top\_p=0.95 Recommended max steps = 300
thinking=
BrowseComp 1.0 per step tokens = 48k;
total max tokens = 256k
1 top\_p=0.95 Recommended max steps = 300
thinking=
DeepSearchQA 1.0 per step tokens = 48k;
total max tokens = 256k
1 top\_p=0.95 Recommended max steps = 300
thinking=
WideSearch 1.0 per step tokens = 48k;
total max tokens = 256k
4 top\_p=0.95 Recommended max steps = 300
thinking=
Toolathlon 1.0 per step tokens = 48k;
total max tokens = 256k
4 top\_p=0.95 Recommended max steps = 300
thinking=
MCPMark 1.0 per step tokens = 48k;
total max tokens = 256k
4 top\_p=0.95 Recommended max steps = 300
thinking=
Claw Eval 1.0 per step tokens = 48k;
total max tokens = 256k
4 top\_p=0.95 Recommended max steps = 300
thinking=
APEX-Agents 1.0 per step tokens = 48k;
total max tokens = 256k
4 top\_p=0.95 Recommended max steps = 300
thinking=
Coding Terminal-Bench 2.0 (Terminus-2) 1.0 max tokens = 256k 3 top\_p=0.95 thinking=
SWE-Bench Pro 1.0 per step tokens = 32k;
total max tokens = 256k
5 top\_p=0.95 Recommended max steps = 300
thinking=
SWE-Bench Multilingual 1.0 per step tokens = 32k;
total max tokens = 256k
5 top\_p=0.95 Recommended max steps = 300
thinking=
SWE-Bench Verified 1.0 per step tokens = 32k;
total max tokens = 256k
5 top\_p=0.95 Recommended max steps = 300
thinking=
SciCode 1.0 max tokens = 96k 4 top\_p=0.95 thinking=
OJBench (python) 1.0 max tokens = 96k 8 top\_p=0.95 thinking=
LiveCodeBench (v6) 1.0 max tokens = 96k 1 top\_p=0.95 thinking=
Math AIME 2026 1.0 max tokens = 96k 32 top\_p=0.95 thinking=
HMMT 2026 (Feb) 1.0 max tokens = 96k 32 top\_p=0.95 thinking=
IMO-AnswerBench 1.0 max tokens = 96k 4 top\_p=0.95 thinking=
Knowledge HLE-Full 1.0 max tokens = 96k 1 top\_p=0.95 thinking=
GPQA-Diamond 1.0 max tokens = 96k 8 top\_p=0.95 thinking=
## K2.5 Models Benchmark Recommended Settings
Benchmark Category Benchmark Temperature Recommended max tokens Recommended runs Top-p Others (e.g. test log)
Multi-modal MMMU-Pro 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
CharXiv (RQ) 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
MathVision 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
MathVista 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
OCRBench 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
ZeroBench 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
WorldVQA 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
InfoVQA (val) 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
SimpleVQA 1.0 max tokens = 64k 3 top\_p=0.95 thinking=
ZeroBench w/ tools 1.0 max tokens = 64k 3 top\_p=0.95 Recommended max steps = 30
thinking=
Code SWE Series 1.0 per step tokens = 16k;
total max tokens = 256k
5 top\_p=0.95 thinking=
Lcb + OJBench 1.0 max tokens = 128k 1 top\_p=0.95 thinking=
TerminalBench 1.0 max tokens = 128k 3 top\_p=0.95 thinking=
Reasoning AIME2025 no tools 1.0 total max tokens = 96k 32 top\_p=0.95 thinking=
AIME2025 w/ tools 1.0 per turn tokens = 96k;
total max tokens = 96k
32 top\_p=0.95 thinking=
Recommended max steps = 120
HLE no tools 1.0 max tokens = 96k 1 top\_p=0.95 thinking=
HLE w/ tools 1.0 total max tokens = 128k;
per step tokens = 48k
1 top\_p=0.95 thinking=
Recommended max steps = 120
HLE heavy 1.0 total max tokens = 128k;
per step tokens = 48k
1 top\_p=0.95 thinking=
Recommended max steps = 200
parallel n=8
HMMT2025 no tools 1.0 max tokens = 96k 32 top\_p=0.95 thinking=
HMMT2025 w/tools 1.0 per step tokens = 96k;
total tokens = 96k
32 top\_p=0.95 thinking=
Recommended max steps = 120
IMO-AnswerBench 1.0 max tokens = 96k 3 top\_p=0.95 thinking=
GPQA-Diamond 1.0 max tokens = 96k 8 top\_p=0.95 thinking=
Agentic Search Task BrowseComp / BrowseComp-ZH / Seal-0 / Frames 1.0 per step tokens = 24k;
total max tokens = 256k
4 top\_p=0.95 thinking=
Recommended max steps = 250
Recommend using a context management mechanism to prevent overly long context and ensure enough tool calls
Include today's date in the system prompt and let the model search when it is uncertain
Agentic Task Tau 1.0 >=16k 4 top\_p=0.95 thinking=
Recommended max steps = 100
For third-party providers, refer to Kimi Vendor Verifier (KVV) to choose high-accuracy services. Details: [https://kimi.com/blog/kimi-vendor-verifier.html](https://kimi.com/blog/kimi-vendor-verifier.html) **Tool Use Compatibility** When using tools, if the thinking parameter is set to `{"type": "enabled"}`, please note the following constraints to ensure model performance: * `tool_choice` can only be set to "auto" or "none" (default is "auto") to avoid conflicts between reasoning content and the specified tool\_choice. Any other value will result in an error; * During multi-step tool calling, you must keep the `reasoning_content` from the assistant message in the current turn's tool call within the context, otherwise an error will be thrown; * The official builtin `$web_search` tool is temporarily incompatible with Kimi K2.5/K2.6 thinking mode, you can choose to disable thinking mode first and then use the `$web_search` tool. You can refer to [Use Thinking Mode](/docs/guide/use-thinking-models) for correct usage of tool calling. ## K2-Thinking Series Models Benchmark Recommended Settings
Category Benchmark Temperature Max token Suggested runs Notes
Code SWE 0.7(recommended)
1.0 (ok)
per step tokens = 16k;
total max token = 256k
5
Lcb + OJBench 1.0 max tokens = 128k 1
TerminalBench 1.0 max tokens = 128k 3
Reasoning AIME2025 no tools 1.0 total max tokens = 96k 32
AIME2025 w/ tools 1.0 per step tokens = 48k;
total max tokens = 128k
16 max steps = 120
HLE no tools 1.0 max tokens = 96k 1
HLE w/ tools 1.0 total max tokens = 128k;
per step tokens = 48k
1 max steps = 120
HLE heavy 1.0 total max tokens = 128k;
per step tokens = 48k
1 max steps = 200
parallel n=8
HMMT2025 no tools 1.0 max tokens = 96k 32
HMMT2025 w/tools 1.0 per step tokens = 96k;
total tokens = 96k
32 max steps = 120
IMO-AnswerBench 1.0 max tokens = 96k 3
GPQA-Diamond 1.0 max tokens = 96k 8
Agentic Search Task BrowseComp/ BrowseComp-ZH/Seal-0/ Frames 1.0 per step tokens = 24k;
total max tokens = 256k
4 max steps = 250
Enable context management to prevent context overflow and ensure enough tool calls.
Include today's date in the system prompt, and tell the model to search when unsure.
Agentic Task Tau 0.0 >=16k 4 max steps = 100
## API Recommendations & Notes * **Use the official API:** some 3rd-party endpoints show noticeable accuracy drift. * Use the recommended models for testing * For K2.6: use **`kimi-k2.6`** for testing * For K2.5: use **`kimi-k2.5`** for testing * For K2 series: use **`kimi-k2-thinking-turbo`** for faster inference * **Must set:** `stream = true` * Non-streaming mode can lead to random mid-connection interruptions that are hard to control. * **Current API default settings:** * Kimi K2.6: * default max\_tokens = 32768 * default thinking = `{"type": "enabled", "keep": null}` * default temperature = 1.0 * default top\_p = 0.95 * default n = 1 * default presence\_penalty = 0.0 * default frequency\_penalty = 0.0 * Kimi K2 Thinking: * default temp = 1.0 * default max token = 64000 * Kimi K2.5: * default max\_tokens = 32768 * default thinking = `{"type": "enabled"}` * default temperature = 1.0 * default top\_p = 0.95 * default n = 1 * default presence\_penalty = 0.0 * default frequency\_penalty = 0.0 * **Timeouts:** * With `stream = false`, `api.moonshot.ai` timeout = **2 hours**, but some ISPs may terminate earlier. * So again we recommend you to set `stream = true` * **Concurrency:** * Keep concurrency low to avoid rate limiting * **Retry logic** is not optional: * handle overloaded * handle unexpected finish reason due to random server issues * handle errors due to complicated network issues ## FAQ **Q1. Is the temperature setting consistent across models?** **A.** No. Different model families use different recommended temperatures: * k2.6 model: temperature = 1.0 * k2.5 model: temperature = 1.0 * k2-thinking series: temperature = 1.0 * k2 other series: temperature = 0.6 **Q2. Why use stream = true?** **A.** Long outputs can take minutes. Idle TCP connections may be terminated by firewalls, load balancers, or NAT gateways. **Streaming keeps the connection alive** and significantly improves reliability. In production, requests with stream = false fail far more often than with stream = true. **Q3. How much concurrency should I use?** **A.** Your API account has specific rate limits (see [Recharge and Rate Limits](/docs/pricing/limits)). Start low. If you hit **HTTP 429** (rate limit), your concurrency is too high. **Accuracy > speed,** so tune concurrency to stay within limits. **Q5. Why should I add retry?** **A.** Even with streaming, requests can fail due to transient network issues. **Retry** on temporary faults (network jitter, server overload, rate limiting) to avoid avoidable failures. **Q6. Why should multi-turn or multi-step tasks include full context and reasoning?** **A.** The model needs full context to stay logically consistent. Without previous reasoning steps, later turns can go off track or produce incomplete answers. ## Contact Us Hit any issues? Drop us an email at [**api-service@moonshot.ai**](mailto:api-service@moonshot.ai) with your logs. We'll take a look! # Use Kimi in Claude Code Source: https://platform.kimi.ai/docs/guide/claude-code-kimi Configure Claude Code with Kimi API environment variables and review model selection, verification steps, and compatibility boundaries. > [Claude Code](https://claude.com/product/claude-code) is a programming agent product from Anthropic. Its interface, configuration options, and supported capabilities may change across versions. This guide describes a general integration approach: forwarding Claude Code model requests to the Kimi API through environment variables. ## Install Claude Code Skip this step if Claude Code is already installed. Run: ```shell theme={null} npm install -g @anthropic-ai/claude-code ``` macOS and Linux: ```shell theme={null} # Install Node.js curl -fsSL https://fnm.vercel.app/install | bash # Open a new terminal so fnm takes effect fnm install 24.3.0 fnm default 24.3.0 fnm use 24.3.0 ``` Windows (PowerShell): ```powershell theme={null} # Right-click the Windows button, click "Terminal", then run: winget install OpenJS.NodeJS Set-ExecutionPolicy -Scope CurrentUser RemoteSigned # Close the terminal window and open a new one ``` After installing Node.js, run the initialization once: ```shell theme={null} node --eval " const fs = require('fs'); const path = require('path'); const os = require('os'); const homeDir = os.homedir(); const filePath = path.join(homeDir, '.claude.json'); if (fs.existsSync(filePath)) { const content = JSON.parse(fs.readFileSync(filePath, 'utf-8')); fs.writeFileSync(filePath, JSON.stringify({ ...content, hasCompletedOnboarding: true }, null, 2), 'utf-8'); } else { fs.writeFileSync(filePath, JSON.stringify({ hasCompletedOnboarding: true }, null, 2), 'utf-8'); }" ``` If you previously modified `~/.claude/settings.json` through third-party tools or by hand, stale entries in its `env` field **override** environment variables exported in your terminal, which can silently prevent the new configuration from taking effect or rewrite model requests. Run the following script to clean them up first: ```shell theme={null} node --eval " const fs = require('fs'); const path = require('path'); const os = require('os'); const settingsPath = path.join(os.homedir(), '.claude', 'settings.json'); if (fs.existsSync(settingsPath)) { const content = JSON.parse(fs.readFileSync(settingsPath, 'utf-8')); if (content && typeof content === 'object' && content.env && typeof content.env === 'object') { for (const key of [ 'ANTHROPIC_BASE_URL', 'ANTHROPIC_API_KEY', 'ANTHROPIC_AUTH_TOKEN', 'ANTHROPIC_MODEL', 'ANTHROPIC_SMALL_FAST_MODEL', 'CLAUDE_CODE_SUBAGENT_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL', 'ANTHROPIC_DEFAULT_OPUS_MODEL_NAME', 'ANTHROPIC_DEFAULT_SONNET_MODEL', 'ANTHROPIC_DEFAULT_SONNET_MODEL_NAME', 'ANTHROPIC_DEFAULT_HAIKU_MODEL', 'ANTHROPIC_DEFAULT_HAIKU_MODEL_NAME', 'ANTHROPIC_DEFAULT_FABLE_MODEL', 'ANTHROPIC_DEFAULT_FABLE_MODEL_NAME', 'ENABLE_TOOL_SEARCH', 'CLAUDE_CODE_AUTO_COMPACT_WINDOW', 'CLAUDE_CODE_EFFORT_LEVEL', ]) { delete content.env[key]; } fs.writeFileSync(settingsPath, JSON.stringify(content, null, 2), 'utf-8'); } }" ``` This script only removes endpoint, credential, and model-related variables from `env`; other settings (such as permissions and theme) are left untouched. Also check shell configuration files such as `~/.zshrc` and `~/.bashrc` for stale `ANTHROPIC_*` exports (on Windows, check your user environment variables) and remove any you find — they interfere with the new configuration just as much. ## Get a Kimi API Key Visit [Kimi Open Platform](https://platform.kimi.ai/console/api-keys) to create an API key (choose the default project), and use it in place of `YOUR_MOONSHOT_API_KEY` below. ## Configure Environment Variables Choose **one of the two methods below and do not mix them**: Method 1 takes effect immediately but only lasts for the current terminal session; Method 2 writes to a configuration file and persists. ### Method 1: Terminal Environment Variables (Current Session Only) macOS and Linux: ```shell theme={null} export ANTHROPIC_BASE_URL="https://api.moonshot.ai/anthropic" export ANTHROPIC_AUTH_TOKEN="${YOUR_MOONSHOT_API_KEY}" export ANTHROPIC_MODEL="kimi-k3[1m]" export ANTHROPIC_DEFAULT_OPUS_MODEL="kimi-k3[1m]" export ANTHROPIC_DEFAULT_SONNET_MODEL="kimi-k3[1m]" export ANTHROPIC_DEFAULT_HAIKU_MODEL="kimi-k3[1m]" export ANTHROPIC_DEFAULT_FABLE_MODEL="kimi-k3[1m]" export CLAUDE_CODE_SUBAGENT_MODEL="kimi-k3[1m]" export CLAUDE_CODE_AUTO_COMPACT_WINDOW="1048576" export CLAUDE_CODE_EFFORT_LEVEL="max" claude ``` Windows (PowerShell): ```powershell theme={null} $env:ANTHROPIC_BASE_URL="https://api.moonshot.ai/anthropic"; $env:ANTHROPIC_AUTH_TOKEN="YOUR_MOONSHOT_API_KEY" $env:ANTHROPIC_MODEL="kimi-k3[1m]" $env:ANTHROPIC_DEFAULT_OPUS_MODEL="kimi-k3[1m]" $env:ANTHROPIC_DEFAULT_SONNET_MODEL="kimi-k3[1m]" $env:ANTHROPIC_DEFAULT_HAIKU_MODEL="kimi-k3[1m]" $env:ANTHROPIC_DEFAULT_FABLE_MODEL="kimi-k3[1m]" $env:CLAUDE_CODE_SUBAGENT_MODEL="kimi-k3[1m]" $env:CLAUDE_CODE_AUTO_COMPACT_WINDOW="1048576" $env:CLAUDE_CODE_EFFORT_LEVEL="max" claude ``` ### Method 2: Write To settings.json (Persistent) Write the same variables into the `env` field of `~/.claude/settings.json`: ```json theme={null} { "env": { "ANTHROPIC_BASE_URL": "https://api.moonshot.ai/anthropic", "ANTHROPIC_AUTH_TOKEN": "YOUR_MOONSHOT_API_KEY", "ANTHROPIC_MODEL": "kimi-k3[1m]", "ANTHROPIC_DEFAULT_OPUS_MODEL": "kimi-k3[1m]", "ANTHROPIC_DEFAULT_SONNET_MODEL": "kimi-k3[1m]", "ANTHROPIC_DEFAULT_HAIKU_MODEL": "kimi-k3[1m]", "ANTHROPIC_DEFAULT_FABLE_MODEL": "kimi-k3[1m]", "CLAUDE_CODE_SUBAGENT_MODEL": "kimi-k3[1m]", "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1048576", "CLAUDE_CODE_EFFORT_LEVEL": "max" } } ``` Note: `env` values in `settings.json` **override** variables exported in your terminal; this file contains your API key in plaintext — do not commit it to a git repository; restart Claude Code after saving. ### Configuration Reference Claude Code uses different model tiers for different scenarios (main conversation, background summarization, sub-agents, and so on). Configuring only some of the variables makes the corresponding scenarios fail silently: | Variable | Purpose | Impact if missing or misconfigured | | ------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `ANTHROPIC_BASE_URL` | Forwards model requests to Kimi's Anthropic-compatible endpoint | Requests go to Anthropic's official endpoint and fail authentication | | `ANTHROPIC_AUTH_TOKEN` | Authenticates with your Kimi API key | Returns 401 authentication errors | | `ANTHROPIC_MODEL` | Model used for the main conversation | Falls back to a default Claude model name that the Kimi endpoint cannot recognize, resulting in model-not-found errors | | `ANTHROPIC_DEFAULT_OPUS_MODEL` / `ANTHROPIC_DEFAULT_SONNET_MODEL` / `ANTHROPIC_DEFAULT_HAIKU_MODEL` / `ANTHROPIC_DEFAULT_FABLE_MODEL` | Model names Claude Code uses when selecting models by task tier | Tasks on the corresponding tier (e.g. background title generation and summarization on the haiku tier) fail | | `CLAUDE_CODE_SUBAGENT_MODEL` | Model used by sub-agents | Sub-agent tasks fail or degrade noticeably | | `CLAUDE_CODE_AUTO_COMPACT_WINDOW` | Context window size that triggers automatic compaction | Must match the model's context: `1048576` for `kimi-k3` (1M), `262144` for `kimi-k2.7-code` (256K). Too small compacts prematurely and loses context; too large causes context-length errors | | `CLAUDE_CODE_EFFORT_LEVEL` | Controls Claude Code's reasoning effort | Set to `max` to enable the most thorough reasoning; lower values may reduce quality for complex tasks | ## Models And Thinking Behavior How the three models actually behave in Claude Code (verified against the Anthropic-compatible endpoint): | Model | Thinking | Notes | | -------------------------------- | ------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `kimi-k3` (default on this page) | On by default | Works out of the box, no extra configuration needed | | `kimi-k2.7-code` | Always on | Requests must explicitly enable thinking — keep Thinking on (press `Tab`) in Claude Code. With thinking off, requests are rejected (`invalid thinking: only type=enabled is allowed for this model`) | | `kimi-k2.6` | Optional | Can run with thinking off; good for latency-sensitive simple tasks | When switching models, replace the value of every model variable in the configuration with the new model name. ## Confirm That The Configuration Took Effect In Claude Code, enter `/status` to confirm the configuration: * Base URL should show `https://api.moonshot.ai/anthropic` * Model should show `kimi-k3[1m]` Claude Code's `/model` menu is a built-in list of fixed aliases — it **does not show Kimi models**, and you do not need to switch anything there. Whether the configuration took effect is determined by what `/status` shows. status Finally, send any message (for example `hi`). Receiving a normal reply confirms the end-to-end setup works. ## Turn On Thinking `kimi-k3` thinks by default and works out of the box. If you switch to `kimi-k2.7-code`, it requires requests to explicitly enable thinking: press `Tab` in Claude Code to turn Thinking on and start working once you see the "Thinking on" indicator — otherwise the model rejects requests (`400 invalid thinking`), and features such as WebSearch do not work. thinking-on You can now use Claude Code for development normally. ## Switch To The High-Speed Model Kimi K2.7 Code offers a high-speed variant, `kimi-k2.7-code-highspeed`, with an output speed about 5-6x that of the regular version. If you prioritize output speed, change the value of every model variable in the configuration to `kimi-k2.7-code-highspeed` (note that it requires explicitly enabled thinking — see Models And Thinking Behavior). See [K2.7 Code pricing](/docs/pricing/chat-k27-code) for details. ## Third-Party Tool: cc-switch Community tools such as cc-switch can switch between multiple provider configurations. These tools are not maintained by Kimi, and their presets may differ from the values recommended on this page. After using them, verify each variable against the Configuration Reference above, and use `/status` to confirm the Base URL and model actually in effect. ## FAQ * **WebSearch fails with `400 invalid thinking: only type=enabled is allowed for this model`**: `kimi-k2.7-code` requires thinking, and WebSearch requests without it are rejected by the platform. Press `Tab` to turn Thinking on first, then retry; if it still fails, switch to `kimi-k2.6` (thinking optional, not subject to this restriction). `kimi-k3` is not affected. This is unrelated to your local configuration or cc-switch. * **WebFetch reports `temporarily unavailable` or returns no fetched content**: the endpoint does not support WebFetch for now — unrelated to your configuration; it will work once the platform adds support. As a workaround, paste the web page content into the chat, or use an MCP scraping tool instead. Check that `ANTHROPIC_AUTH_TOKEN` is a valid Kimi API key. If you previously configured `ANTHROPIC_API_KEY`, remove it to avoid conflicts with `ANTHROPIC_AUTH_TOKEN` when both are present. Check the spelling of every model variable (`kimi-k3[1m]`), and make sure there are no extra spaces or quotes. This usually means `ANTHROPIC_DEFAULT_HAIKU_MODEL`, `ANTHROPIC_DEFAULT_FABLE_MODEL`, or `CLAUDE_CODE_SUBAGENT_MODEL` is not set, so those scenarios requested a model name the Kimi endpoint cannot recognize. Fill in the missing variables per the Configuration Reference above. * Check for stale entries in the `env` field of `~/.claude/settings.json` (they override terminal environment variables); you can run the cleanup script in the collapsed section above; * Variables exported in a terminal only apply to the current session and must be set again in a new window. If you appended them to `~/.zshrc` or used the settings.json method, make sure you restarted Claude Code after the change. Make sure `ANTHROPIC_BASE_URL` matches the platform where you created the API key — create the key on the platform linked in "Get a Kimi API Key" above and use the endpoint shown on this page. Once set, `ANTHROPIC_AUTH_TOKEN` takes precedence over a saved claude.ai login, so no action is usually needed. Run `/status` in a session to confirm which credential source is active; to remove the saved login, run `/logout`. # Use Kimi K3 in Codex CLI Source: https://platform.kimi.ai/docs/guide/codex-kimi Connect Codex CLI to Kimi Open Platform with CC Switch, configure `kimi-k3`, and verify text and visual inputs. This guide explains how to connect Codex CLI to Kimi Open Platform through CC Switch and use the `kimi-k3` model. Codex CLI currently supports text and image input, but does not provide a native video input channel — you cannot submit a video file directly as multimodal input to the model. To analyze video using only Codex CLI's existing input methods, you can first extract key frames with ffmpeg and optionally combine them with audio transcription before handing them to the model. Note that this is a limitation of Codex CLI's input layer, not of the Kimi K3 model — the Kimi K3 API natively supports video input. Call `kimi-k3` directly as described in [Vision Input](/docs/guide/use-kimi-vision-model) for full video understanding, with no manual frame extraction required. ## Prerequisites Complete the following preparations first. Follow the corresponding official instructions for installation and account-related operations; this guide does not repeat those procedures. Follow the official Codex documentation and start Codex CLI at least once. Create and save an API key in Kimi Open Platform. Follow the official CC Switch instructions to install the version for your operating system. CC Switch is a third-party open-source tool and is not part of Kimi Open Platform. Before using it, evaluate it according to your organization's security and compliance requirements; your API key and Codex requests and responses will be processed by its local router. After installation, open CC Switch and follow the steps below. ## Step 1: Enable Codex Routing In CC Switch, open **Settings > Routing**, then: 1. Turn on **Routing Master Switch** to start the local routing service. 2. Under **Routing Enabled**, turn on **Codex**. Enable CC Switch Local Routing and Codex routing Codex CLI uses the Responses API, while Kimi Open Platform provides an OpenAI-compatible Chat Completions API. CC Switch Local Routing converts requests and streaming responses between the two protocols. Keep CC Switch and Codex routing running while using the Kimi provider. ## Step 2: Add the Kimi Provider 1. Return to the CC Switch home screen and select the **Codex** tab at the top. 2. Click **+** in the upper-right corner to add a provider. Open the Codex tab and add a provider 3. Confirm that you are on the **Codex Provider** page. Search for Kimi if needed, then select **Kimi** from the preset provider list. Select the Kimi provider preset 4. Enter the following settings: | Setting | Value | | -------------------------- | ----------------------------------------- | | API request URL (Base URL) | `https://api.moonshot.ai/v1` | | API key | The API key created in Kimi Open Platform | | Default model | `kimi-k3` | 5. After changing the default model to `kimi-k3`, click **Add to mapping**. Enter the Kimi API key, request URL, and default model 6. Scroll down and confirm the following advanced settings: | Setting | Value | | ------------------------- | ------------------------------------- | | Upstream format | `Chat Completions (routing required)` | | Prompt cache routing | `Auto (recommended)` | | Supports thinking mode | On | | Supports reasoning effort | On | | Menu display name | `kimi-k3` | | Actual request model | `kimi-k3` | | Context window | `1048576` | 7. After confirming the settings, click **Add** in the lower-right corner. Configure the Kimi upstream format, reasoning capabilities, and model mapping ## Step 3: Enable the Kimi Provider Return to the Codex provider list after adding the provider and click **Enable** on the Kimi provider you just added. Enable the Kimi provider Confirm that: * Kimi is the active provider under the Codex tab; * CC Switch Local Routing is running; * Codex is enabled under the routing settings. ## Step 4: Start Codex CLI If Codex CLI is already running, exit the current session. Enter the project directory where you want to work and start Codex CLI again: ```bash theme={null} cd /path/to/your/project codex ``` Restarting allows Codex CLI to load the latest provider and model configuration written by CC Switch. After startup, confirm that Codex CLI shows `kimi-k3` as the current model, then send a simple request: ```text theme={null} hello ``` If Codex CLI returns a response and the status bar shows `kimi-k3`, the configuration is working. Verify kimi-k3 in Codex CLI You can also check the CC Switch routing counter or request logs for a new Codex request. # Configure ModelScope MCP Server in Playground Source: https://platform.kimi.ai/docs/guide/configure-the-modelscope-mcp-server Sync and enable ModelScope-hosted MCP services in Kimi Playground so models can call your configured tools. Through an official partnership between the Kimi API Platform and ModelScope, you can sync all hosted MCP service configurations under your ModelScope account into Kimi Playground with a single API token. Follow this page when you want models in Playground to call MCP tools. ## Sync ModelScope-Hosted MCP Services Log in to Kimi Playground ([https://platform.kimi.ai/playground](https://platform.kimi.ai/playground)) and make sure you can have basic conversations with the Kimi K2 model. MCP services are added in "MCP Server Settings", where ModelScope is selected as the default MCP service provider. If you haven't used the ModelScope MCP marketplace before, first refer to the [ModelScope official documentation](https://modelscope.cn/mcp/kimi-playground) to select and host your MCP services; you can also discover numerous MCP servers in the ModelScope community. ### Open MCP Server Settings Click the configuration button to open "MCP Server Settings": mcp-server-setting ### Enter Your ModelScope API Token and Sync In the panel that appears, choose to sync with the external platform: syc You can obtain the API token from the [ModelScope Homepage - Access Token](https://modelscope.cn/my/myaccesstoken) page: keys Paste the token into the field in Step 3 and click "Start Sync": start-syc Once syncing completes, all configured and connected ModelScope Hosted MCP services appear in the available MCP services list in Kimi Playground: mcp-list ### Incrementally Sync MCP Services If you later add or remove hosted MCP services in the ModelScope MCP marketplace, click the sync button in "Settings - MCP Server - Sync Server" to perform an incremental update: add-mcp ## Enable MCP Services in a Conversation After syncing, the imported "MCP Services List" appears on the left side of the Kimi Playground page. Multi-select and enable the MCP services you want to use in the current conversation: manage-mcp # Set Parameters for Multi-turn Chat Source: https://platform.kimi.ai/docs/guide/engage-in-multi-turn-conversations-using-kimi-api Maintain message history, control context length, and construct multi-turn conversations with the stateless Kimi API. Unlike the Kimi intelligent assistant, the Kimi API is **stateless** and has no memory of its own: across multiple requests, the model doesn't know what you asked in a previous request and won't remember any context—if you tell it you are 27 years old in one request, it won't know that in the next. To enable multi-turn conversations, manually maintain the context for each request by sending the conversation history along with the next request, so the model can see what has been discussed before. The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ## Give the model memory with the messages list The following example modifies the one from the previous chapter to show how maintaining a `messages` list gives the model memory: each turn appends both the user's new message (role=user) and the model's reply (role=assistant) to the list, then sends the whole list with the request. The key points are annotated as comments in the code: ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key = os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url = "https://api.moonshot.ai/v1", ) # We define a global variable messages to keep track of the historical conversation messages between us and the Kimi large language model # The messages include both the questions we ask the Kimi large language model (role=user) and the replies it gives us (role=assistant) # Of course, it also includes the initial System Prompt (role=system) # The messages in the list are arranged in chronological order messages = [ {"role": "system", "content": "You are Kimi, an artificial intelligence assistant provided by Moonshot AI. You are better at conversing in Chinese and English. You provide users with safe, helpful, and accurate answers. At the same time, you refuse to answer any questions involving terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages."}, ] def chat(input: str) -> str: """ The chat function supports multi-turn conversations. Each time the chat function is called to converse with the Kimi large language model, the model will 'see' the historical conversation messages that have already been generated. In other words, the Kimi large language model has a memory. """ global messages # We construct the user's latest question as a message (role=user) and add it to the end of the messages list messages.append({ "role": "user", "content": input, }) # We converse with the Kimi large language model, carrying the messages along completion = client.chat.completions.create( model="kimi-k3", messages=messages ) # Through the API, we receive the reply message (role=assistant) from the Kimi large language model assistant_message = completion.choices[0].message # To give the Kimi large language model a complete memory, we must also add the message it returns to us to the messages list messages.append(assistant_message) return assistant_message.content print(chat("Hello, I am 27 years old this year.")) print(chat("Do you know how old I am this year?")) # Here, based on the previous context, the Kimi large language model will know that you are 27 years old ``` ```js theme={null} const OpenAI = require("openai") const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }); // We define a global variable messages to keep track of the historical conversation messages between us and the Kimi large language model // The messages include both the questions we ask the Kimi large language model (role=user) and the replies it gives us (role=assistant) // Of course, it also includes the initial System Prompt (role=system) // The messages in the list are arranged in chronological order let messages = [ { role: "system", content: "You are Kimi, an artificial intelligence assistant provided by Moonshot AI. You are better at conversing in Chinese and English. You provide users with safe, helpful, and accurate answers. At the same time, you refuse to answer any questions involving terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages.", }, ]; async function chat(input) { /** * The chat function supports multi-turn conversations. Each time the chat function is called to converse with the Kimi large language model, the model will 'see' the historical conversation messages that have already been generated. In other words, the Kimi large language model has a memory. */ // We construct the user's latest question as a message (role=user) and add it to the end of the messages list messages.push({ role: "user", content: input, }); // We converse with the Kimi large language model, carrying the messages along const completion = await client.chat.completions.create({ model: "kimi-k3", messages: messages }); // Through the API, we receive the reply message (role=assistant) from the Kimi large language model const assistantMessage = completion.choices[0].message; // To give the Kimi large language model a complete memory, we must also add the message it returns to us to the messages list messages.push(assistantMessage); return assistantMessage.content; } // Example usage (async () => { console.log(await chat("Hello, I am 27 years old this year.")); console.log(await chat("Do you know how old I am this year?")); // Here, based on the previous context, the Kimi large language model will know that you are 27 years old })(); ``` Key points: * The Kimi API has no built-in context memory; use the `messages` parameter to manually tell the model what has been discussed before; * The `messages` list must store both the user's questions (role=user) and the model's replies (role=assistant). ## Truncate history to control context length As the number of `chat` calls grows, the `messages` list keeps getting longer, so each request consumes more Tokens—eventually the messages in the list will exceed the context window supported by the model. Use a strategy to keep the `messages` list within a manageable range, for example by keeping only the latest 20 messages as the context for each request. The following example shows how the `make_messages` function controls the number of messages in each request (keeping the latest 20 by default)—note how it ensures the System Messages remain in the list even when truncating: ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key = os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url = "https://api.moonshot.ai/v1", ) # We place the System Messages in a separate list because every request should carry the System Messages. system_messages = [ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are more proficient in conversing in Chinese and English. You provide users with safe, helpful, and accurate responses. You also reject any questions involving terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages."}, ] # We define a global variable messages to record the historical conversation messages between us and the Kimi large language model. # The messages include both the questions we pose to the Kimi large language model (role=user) and the replies from the Kimi large language model (role=assistant). # The messages are arranged in chronological order. messages = [] def make_messages(input: str, n: int = 20) -> list[dict]: """ The make_messages function controls the number of messages in each request to keep it within a reasonable range, such as the default value of 20. When building the message list, we first add the System Prompt because it is essential no matter how the messages are truncated. Then, we obtain the latest n messages from the historical records as the messages for the request. In most scenarios, this ensures that the number of Tokens occupied by the request messages does not exceed the model's context window. """ global messages # First, we construct the user's latest question into a message (role=user) and add it to the end of the messages list. messages.append({ "role": "user", "content": input, }) # new_messages is the list of messages we will use for the next request. Let's build it now. new_messages = [] # Every request must carry the System Messages, so we need to add the system_messages to the message list first. # Note that even if the messages are truncated, the System Messages should still be in the messages list. new_messages.extend(system_messages) # Here, when the historical messages exceed n, we only keep the latest n messages. if len(messages) > n: messages = messages[-n:] new_messages.extend(messages) return new_messages def chat(input: str) -> str: """ The chat function supports multi-turn conversations. Each time the chat function is called to converse with the Kimi large language model, the model can "see" the historical conversation messages that have already been generated. In other words, the Kimi large language model has memory. """ # We converse with the Kimi large language model carrying the messages. completion = client.chat.completions.create( model="kimi-k3", messages=make_messages(input) ) # Through the API, we obtain the reply message from the Kimi large language model (role=assistant). assistant_message = completion.choices[0].message # To ensure the Kimi large language model has a complete memory, we must add the message returned by the model to the messages list. messages.append(assistant_message) return assistant_message.content print(chat("Hello, I am 27 years old this year.")) print(chat("Do you know how old I am this year?")) # Here, based on the previous context, the Kimi large language model will know that you are 27 years old this year. ``` ```js theme={null} const OpenAI = require("openai") const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }); // We place the System Messages in a separate list because every request should carry the System Messages. const systemMessages = [ { role: "system", content: "You are Kimi, an AI assistant provided by Moonshot AI. You are more proficient in conversing in Chinese and English. You provide users with safe, helpful, and accurate responses. You also reject any questions involving terrorism, racism, pornography, or violence. Moonshot AI is a proper noun and should not be translated into other languages.", }, ]; // We define a global variable messages to record the historical conversation messages between us and the Kimi large language model. // The messages include both the questions we pose to the Kimi large language model (role=user) and the replies from the Kimi large language model (role=assistant). // The messages are arranged in chronological order. let messages = []; async function makeMessages(input, n = 20) { /** * The makeMessages function controls the number of messages in each request to keep it within a reasonable range, such as the default value of 20. When building the message list, we first add the System Prompt because it is essential no matter how the messages are truncated. Then, we obtain the latest n messages from the historical records as the messages for the request. In most scenarios, this ensures that the number of Tokens occupied by the request messages does not exceed the model's context window. */ // First, we construct the user's latest question into a message (role=user) and add it to the end of the messages list. messages.push({ role: "user", content: input, }); // newMessages is the list of messages we will use for the next request. Let's build it now. let newMessages = []; // Every request must carry the System Messages, so we need to add the systemMessages to the message list first. // Note that even if the messages are truncated, the System Messages should still be in the messages list. newMessages = systemMessages.concat(newMessages); // Here, when the historical messages exceed n, we only keep the latest n messages. if (messages.length > n) { messages = messages.slice(-n); } newMessages = newMessages.concat(messages); return newMessages; } async function chat(input) { /** * The chat function supports multi-turn conversations. Each time the chat function is called to converse with the Kimi large language model, the model can "see" the historical conversation messages that have already been generated. In other words, the Kimi large language model has memory. */ // We converse with the Kimi large language model carrying the messages. const completion = await client.chat.completions.create({ model: "kimi-k3", messages: await makeMessages(input) }); // Through the API, we obtain the reply message from the Kimi large language model (role=assistant). const assistantMessage = completion.choices[0].message; // To ensure the Kimi large language model has a complete memory, we must add the message returned by the model to the messages list. messages.push(assistantMessage); return assistantMessage.content; } (async () => { console.log(await chat("Hello, I am 27 years old this year.")); console.log(await chat("Do you know how old I am this year?")); // Here, based on the previous context, the Kimi large language model will know that you are 27 years old this year. })(); ``` ## What else to consider in production The examples above only cover the simplest invocation scenario. In real business logic, you may need to handle more scenarios and edge cases: * In concurrent scenarios, additional read-write locks may be needed; * For multi-user scenarios, maintain a separate `messages` list for each user; * Persist the `messages` list; * Use a more precise way to determine how many messages to retain in the `messages` list; * Summarize the discarded messages and add the summary as a new message to the `messages` list; * …… # Use Kimi API Platform with Kimi Code CLI Source: https://platform.kimi.ai/docs/guide/kimi-code-cli Connect, switch, and update Kimi Code CLI with a Kimi API Platform API Key. [Kimi Code](https://www.kimi.com/code/) is an intelligent coding service for developers, and Kimi Code CLI is its terminal-based AI agent. Kimi Code CLI can directly use an API Key created on [Kimi API Platform](https://platform.kimi.ai/). This guide explains how to: * Connect with an API Key from `platform.kimi.ai` * Switch an existing Kimi Code CLI setup to Kimi API Platform * Replace a configured API Key This guide assumes that Kimi Code CLI is already installed. If it is not installed, see the [official Kimi Code CLI installation guide](https://www.kimi.com/code/docs/en/kimi-code-cli/guides/getting-started.html), or select your system and run: ```bash theme={null} curl -fsSL https://code.kimi.com/kimi-code/install.sh | bash ``` ```powershell theme={null} irm https://code.kimi.com/kimi-code/install.ps1 | iex ``` Before first launch on Windows, install [Git for Windows](https://gitforwindows.org/). Kimi Code CLI uses its bundled Git Bash as the Shell environment. If Git Bash is installed in a non-standard location, set `KIMI_SHELL_PATH` to the absolute path of `bash.exe`. The script automatically downloads the latest release, verifies its checksum, and places the `kimi` executable on your `PATH`. ## Prepare an API Key Open [Kimi API Platform](https://platform.kimi.ai/), sign in, and go to the [API Keys](https://platform.kimi.ai/console/api-keys) page. Create and copy an API Key. Keep your API Key secure. Do not share it or expose the complete value in screenshots. ## Connect Kimi API Platform ### 1. Start Kimi Code CLI Open the project directory where you want to use Kimi Code CLI, then start it: ```bash theme={null} kimi ``` In the interactive interface, enter: ```text theme={null} /login ``` Kimi Code CLI opens the platform selector. Enter /login in Kimi Code CLI ### 2. Select the API Key platform Select the option that matches where the API Key was created: | API Key source | Select in Kimi Code CLI | | ------------------ | -------------------------------------------- | | `platform.kimi.ai` | `Kimi Platform (API key · platform.kimi.ai)` | Use the arrow keys to select the platform, then press `Enter`. The selected platform must match the site where the API Key was created, or API Key validation will fail. Select the Kimi Platform option for platform.kimi.ai ### 3. Enter the API Key Paste the API Key when prompted, then press `Enter`. Kimi Code CLI validates the API Key and loads the models available to the account. Enter a Kimi API Platform API Key ### 4. Select a model After the API Key is validated, Kimi Code CLI displays the models available to the account. Select the model you want to use and confirm. Kimi Code CLI then: * Switches the current session to the selected model * Saves the selected platform and model * Reuses this configuration on later launches When the configuration confirmation appears, Kimi API Platform is connected. Select a Kimi API Platform model ### 5. Verify the connection Run the following command to view the current session status: ```text theme={null} /status ``` Confirm the current model, then send a simple task, for example: ```text theme={null} Review the current project and briefly describe its directory structure. ``` If Kimi Code CLI returns a normal response, the connection is working. Kimi Code CLI responding through Kimi API Platform ## Switch to a Kimi API Platform API Key If Kimi Code CLI already uses another login method, you do not need to exit the program or edit the configuration manually. Run the following command again in the current session: ```text theme={null} /login ``` Then: 1. Select `Kimi Platform (API key · platform.kimi.ai)` 2. Enter the Kimi API Platform API Key 3. Select the model you want to use 4. Wait for the configuration confirmation The current session switches directly to the selected Kimi API Platform model. You do not need to restart Kimi Code CLI. ## Replace the API Key To replace a configured API Key, run `/login` again, select `Kimi Platform (API key · platform.kimi.ai)`, and enter the new API Key. After configuration completes, Kimi Code CLI uses the new API Key. ## Troubleshooting ### API Key validation fails Check the following: * The selected platform is where the API Key was created * The complete API Key was copied without extra spaces * The API Key is still valid * The corresponding Kimi API Platform account can call the API If the API Key was revoked or exposed, create a new API Key in Kimi API Platform and run `/login` again. ### `/login` is unavailable `/login` must be run while Kimi Code CLI is idle. If it is generating content or running a task, wait for the task to finish, or press `Esc` or `Ctrl-C` to interrupt it before running `/login` again. ### No models are displayed Confirm that the API Key and selected platform match, then run `/login` again. If the models still do not load, upgrade Kimi Code CLI and check the Kimi API Platform account status. ## Learn more * [Get started with Kimi Code CLI](https://www.kimi.com/code/docs/en/kimi-code-cli/guides/getting-started.html) * [Kimi Code CLI slash commands](https://www.kimi.com/code/docs/en/kimi-code-cli/reference/slash-commands.html) * [Kimi Code CLI providers and models](https://www.kimi.com/code/docs/en/kimi-code-cli/configuration/providers.html) * [Kimi Code CLI configuration files](https://www.kimi.com/code/docs/en/kimi-code-cli/configuration/config-files.html) * [Kimi API Platform quickstart](/docs/overview) # Kimi K2.6 Source: https://platform.kimi.ai/docs/guide/kimi-k2-6-quickstart Explore Kimi K2.6 text, image, and video understanding, thinking mode, tool calling, and its 256K-token context window. ## Overview of Kimi K2.6 Model Kimi K2.6 is Kimi's general-purpose model, possessing stronger and more stable long-term code writing capabilities, significantly improved instruction compliance and self-correction capabilities, and the ability to handle more complex software engineering tasks. It also significantly enhances the autonomous execution capabilities of the Agent. It supports text, image, and video input, thinking and non-thinking modes, and dialogue and Agent tasks. [Tech Blog](https://www.kimi.com/blog/kimi-k2-6) kimi-k2.6 ### Long-horizon coding capability breakthrough * K2.6 has achieved a breakthrough in long-horizon coding tasks, demonstrating more reliable generalization across diverse programming languages (such as Rust, Go, and Python) and task scenarios (including frontend development, DevOps, and performance optimization). ### Ultra-Long Context Support * `kimi-k2.6`, `kimi-k2.5`, `kimi-k2-0905-preview`, `kimi-k2-turbo-preview`, `kimi-k2-thinking`, and `kimi-k2-thinking-turbo` models all provide a 256K context window. ### Long-Thinking Capabilities * Kimi K2.6 still has strong reasoning capabilities, supporting multi-step tool invocation and reasoning, excelling at solving complex problems, such as complex logical reasoning, mathematical problems, and code writing. ## Example Usage Here is a complete usage example to help you quickly get started with the Kimi K2.6 model. ### Install the OpenAI SDK Kimi API is fully compatible with OpenAI's API format. You can install the OpenAI SDK as follows: ```bash theme={null} pip install --upgrade 'openai>=1.0' ``` ### Verify the Installation ```bash theme={null} python -c 'import openai; print("version =",openai.__version__)' # The output may be version = 1.10.0, indicating the OpenAI SDK was installed successfully and your Python environment is using OpenAI SDK v1.10.0. ``` ## Quick Start * [Try it now](https://platform.kimi.ai/playground): Test model performance in your business scenarios through interactive operations in the Dev Workbench * [Apply for API Key](https://platform.kimi.ai/console/api-keys): Test via API call immediately ### Image Understanding Code Example ```python theme={null} import os import base64 from openai import OpenAI client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1", ) # Replace kimi.png with the path to the image you want Kimi to analyze image_path = "kimi.png" with open(image_path, "rb") as f: image_data = f.read() # Use the standard library base64.b64encode function to encode the image into base64 format image_url = f"data:image/{os.path.splitext(image_path)[1].lstrip('.')};base64,{base64.b64encode(image_data).decode('utf-8')}" completion = client.chat.completions.create( model="kimi-k2.6", messages=[ {"role": "system", "content": "You are Kimi."}, { "role": "user", # Note: content is changed from str type to a list containing multiple content parts. # Image (image_url) is one part, and text is another part. "content": [ { "type": "image_url", # <-- Use image_url type to upload images, with content as base64-encoded image data "image_url": { "url": image_url, }, }, { "type": "text", "text": "Please describe the content of the image.", # <-- Use text type to provide text instructions }, ], }, ], ) print(completion.choices[0].message.content) ``` If your code runs successfully with no errors, you will see output similar to the following: ```text theme={null} [Image description output] ``` ### Video Understanding Code Example ```python theme={null} import os import base64 from openai import OpenAI client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1", ) # Replace kimi.mp4 with the path to the video you want Kimi to analyze video_path = "kimi.mp4" with open(video_path, "rb") as f: video_data = f.read() # Use the standard library base64.b64encode function to encode the video into base64 format video_url = f"data:video/{os.path.splitext(video_path)[1].lstrip('.')};base64,{base64.b64encode(video_data).decode('utf-8')}" completion = client.chat.completions.create( model="kimi-k2.6", messages=[ {"role": "system", "content": "You are Kimi."}, { "role": "user", # Note: content is changed from str type to a list containing multiple content parts. # Video (video_url) is one part, and text is another part. "content": [ { "type": "video_url", # <-- Use video_url type to upload videos, with content as base64-encoded video data "video_url": { "url": video_url, }, }, { "type": "text", "text": "Please describe the content of the video.", # <-- Use text type to provide text instructions }, ], }, ], ) print(completion.choices[0].message.content) ``` ### Multimodal Tool Capability Example Kimi K2.6 model combines multiple capabilities. The following example demonstrates K2.6's visual understanding + tool calling capabilities. First, download this sample video to your local machine, such as `~/Download/test_video.mp4` Then run the following code: ```python theme={null} import base64 import json import os import subprocess import tempfile from pathlib import Path from openai import OpenAI tools = [{ "type": "function", "function": { "name": "watch_video_clip", "description": "Watch a video file or a sub-clip of it. If start_time and end_time are not provided, the entire video will be returned.", "parameters": { "type": "object", "properties": { "path": { "type": "string", "description": "The path to the video file to watch" }, "start_time": { "type": "number", "description": "The start time of the clip in seconds (optional, defaults to 0)" }, "end_time": { "type": "number", "description": "The end time of the clip in seconds (optional, defaults to end of video)" } }, "required": ["path"] } } }] def watch_video_clip(path: str, start_time: float | None = None, end_time: float | None = None) -> list[dict]: """ Watch a video file or a sub-clip of it. Args: path: The path to the video file to watch start_time: The start time in seconds (optional, defaults to 0) end_time: The end time in seconds (optional, defaults to end of video) Returns: A list of content blocks in MultiModal Tool API format """ video_path = Path(path) if not video_path.exists(): raise FileNotFoundError(f"Video file not found: {path}") # Get video duration if needed if start_time is None and end_time is None: # Return entire video with open(path, "rb") as f: video_base64 = base64.b64encode(f.read()).decode("utf-8") return [ {"type": "video_url", "video_url": {"url": f"data:video/mp4;base64,{video_base64}"}}, {"type": "text", "text": f"Full video: {video_path.name}"} ] # Get video duration for defaults probe = subprocess.run( ["ffprobe", "-v", "quiet", "-print_format", "json", "-show_format", path], capture_output=True, text=True ) duration = float(json.loads(probe.stdout)["format"]["duration"]) start_time = start_time or 0 end_time = end_time or duration clip_duration = end_time - start_time # Extract clip with tempfile.NamedTemporaryFile(suffix=".mp4", delete=False) as tmp: tmp_path = tmp.name try: subprocess.run([ "ffmpeg", "-y", "-ss", str(start_time), "-i", path, "-t", str(clip_duration), "-c:v", "libx264", "-c:a", "aac", "-preset", "fast", "-crf", "23", "-movflags", "+faststart", "-loglevel", "error", tmp_path ], check=True) with open(tmp_path, "rb") as f: video_base64 = base64.b64encode(f.read()).decode("utf-8") return [ {"type": "video_url", "video_url": {"url": f"data:video/mp4;base64,{video_base64}"}}, {"type": "text", "text": f"Clip from {video_path.name}: {start_time}s - {end_time}s"} ] finally: if os.path.exists(tmp_path): os.unlink(tmp_path) client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1" ) def agent_loop(user_message: str): """Simple agent loop with multimodal tool support.""" messages = [ {"role": "system", "content": "You are a video analysis assistant. Use watch_video_clip to examine specific portions of videos."}, {"role": "user", "content": user_message} ] while True: response = client.chat.completions.create( model="kimi-k2.6", messages=messages, tools=tools, tool_choice="auto" ) message = response.choices[0].message messages.append(message.model_dump()) # No tool calls = done if not message.tool_calls: return message.content # Execute tool calls for tool_call in message.tool_calls: if tool_call.function.name == "watch_video_clip": args = json.loads(tool_call.function.arguments) result = watch_video_clip( path=args["path"], start_time=args.get("start_time"), end_time=args.get("end_time") ) # Multimodal tool result messages.append({ "role": "tool", "tool_call_id": tool_call.id, "content": result }) # Usage answer = agent_loop("Analyze what happens between seconds 8-13 in ~/Download/test_video.mp4") print(answer) ``` ## Best Practices ### Supported Formats Images are supported in formats: png, jpeg, webp, gif.\ Videos are supported in formats: mp4, mpeg, mov, avi, x-flv, mpg, webm, wmv, 3gpp. ### Token Calculation and Billing Image and video token usage is dynamically calculated. You can use the [token estimation API](/docs/api/estimate) to check the expected token consumption for a request containing images or video before processing. Generally, the higher the resolution of an image, the more tokens it will consume. For videos, the number of tokens depends on the number of keyframes and their resolution—the more keyframes and the higher their resolution, the greater the token consumption. The Vision model uses the same billing method as the `moonshot-v1` model series, with charges based on the total number of tokens processed. For more information, see: For token pricing details, refer to [Model Pricing](/docs/pricing/chat-k26). ### Recommended Resolution We recommend that image resolution should not exceed 4k (4096×2160), and video resolution should not exceed FHD (1920×1080). Higher resolutions will only increase processing time and will not improve the model’s understanding. ### Upload File or Base64? Due to the limitation on the overall size of the request body, for very large videos you **must** use the file upload method to utilize vision capabilities.For images or videos that will be referenced multiple times, it is recommended to use the file upload method. Regarding file upload limitations, please refer to the [File Upload documentation](/docs/api/files-upload). Image quantity limit: The Vision model has no limit on the number of images, but ensure that the request body size does not exceed 100M URL-formatted images: Not supported, currently only supports base64-encoded image content ## Parameters Differences in Request Body Parameters are listed in [chat](/docs/api/chat). However, behaviour of some parameters may be different in k2.6/k2.5 models. **We recommend using the default values instead of manually configuring these parameters.** Differences are listed below. | Field | Required | Description | Type | Values | | ------------------ | -------- | ---------------------------------------------------------------------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------- | | max\_tokens | optional | The maximum number of tokens to generate for the chat completion. | int | Default to be 32k aka 32768 | | thinking | optional | **New!** This parameter controls if the thinking is enabled for this request | object | Default to be `{"type": "enabled"}`. Value can only be one of `{"type": "enabled"}` or `{"type": "disabled"}` | | temperature | optional | The sampling temperature to use | float | k2.6/k2.5 model will use a fixed value 1.0, non-thinking mode will use a fixed value 0.6. Any other value will result in an error | | top\_p | optional | A sampling method | float | k2.6/k2.5 model will use a fixed value 0.95. Any other value will result in an error | | n | optional | The number of results to generate for each input message | int | k2.6/k2.5 model will use a fixed value 1. Any other value will result in an error | | presence\_penalty | optional | Penalizing new tokens based on whether they appear in the text | float | k2.6/k2.5 model will use a fixed value 0.0. Any other value will result in an error | | frequency\_penalty | optional | Penalizing new tokens based on their existing frequency in the text | float | k2.6/k2.5 model will use a fixed value 0.0. Any other value will result in an error | ## Tool Use Compatibility When using tools, if the thinking parameter is set to `{"type": "enabled"}`, please note the following constraints to ensure model performance: * `tool_choice` can only be set to "auto" or "none" (default is "auto") to avoid conflicts between reasoning content and the specified tool\_choice. Any other value will result in an error; * During multi-step tool calling, you must keep the `reasoning_content` from the assistant message in the current turn's tool call within the context, otherwise an error will be thrown; * The official builtin `$web_search` tool is temporarily incompatible with Kimi K2.6/Kimi K2.5 thinking mode, you can choose to disable thinking mode first and then use the `$web_search` tool. You can refer to [Use Thinking Mode](/docs/guide/use-thinking-models) for correct usage of tool calling. ### Disable Thinking Capability Example For the `kimi-k2.6`, `kimi-k2.5` model, you can disable thinking by specifying `"thinking": {"type": "disabled"}` in the request body: ```bash theme={null} $ curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k2.6", "messages": [ {"role": "user", "content": "hello"} ], "thinking": {"type": "disabled"} }' ``` ```python theme={null} import os import openai client = openai.Client( base_url="https://api.moonshot.ai/v1", api_key=os.getenv("MOONSHOT_API_KEY"), ) response = client.chat.completions.create( model="kimi-k2.6", messages=[ {"role": "user", "content": "hello"} ], extra_body={ "thinking": {"type": "disabled"} }, # Pass additional request body via extra_body parameter to disable thinking max_tokens=1024*32 # No need to set temperature ) print(response.choices[0].message.content) print(response) ``` ## Model Pricing For token pricing details, refer to [Model Pricing](/docs/pricing/chat-k26). ## Learn More * For the benchmark testing with Kimi K2.6, please refer to this [benchmark best practice](/docs/guide/benchmark-best-practice) * For the most detailed API usage example of Kimi K2.6, see: [Vision Input](/docs/guide/use-kimi-vision-model) * See how to use Kimi K2 in [Claude Code](/docs/guide/claude-code-kimi) * Learn how to configure and use the [Thinking Mode](/docs/guide/use-thinking-models) The web search (`web_search`) is currently being updated. We do not recommend using this functionality in the near term. This documentation is outdated; please follow subsequent content updates. * Web search is a powerful official tool provided by the Kimi API. See how to use [Web Search](/docs/guide/use-web-search) and other [official tools](/docs/guide/use-official-tools). * For all model pricing see [here](/docs/pricing/chat), [Billing & Rate Limit details](/docs/pricing/limits), and [Web Search Pricing](/docs/pricing/tools) # Kimi K2.7 Code Source: https://platform.kimi.ai/docs/guide/kimi-k2-7-code-quickstart Explore Kimi K2.7 Code and its high-speed variant for coding, multimodal input, thinking, tool calling, and 256K-token contexts. ## Overview of Kimi K2.7 Code Model Kimi K2.7 Code is Kimi's dedicated coding model. It follows instructions more reliably in long contexts, completes coding tasks with higher success rates. External benchmark evaluations show that Kimi K2.7 Code significantly improves instruction compliance and long-horizon coding performance compared to K2.6, while reducing overthinking tendencies by 30% on average. Kimi K2.7 Code HighSpeed (kimi-k2.7-code-highspeed) is the high-speed version of Kimi K2.7 Code, the same model as Kimi K2.7 Code, but with an output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience. (Currently, the resource is limited, and the experience of the high-speed model may be slightly fluctuate,we are gradually increasing the resource.) ### Long-horizon coding capability breakthrough In external benchmark evaluations, Kimi K2.7 Code significantly improves instruction compliance and long-horizon coding performance compared to K2.6, while reducing overthinking tendencies by 30% on average. kimi-k2.7 ### Enhanced Agentic capabilities In external benchmark evaluations, Kimi K2.7 Code significantly improves Agentic capabilities compared to K2.6, with performance improvement of 10%. kimi-k2.7 ### Ultra-Long Context Support * `kimi-k2.7-code`, `kimi-k2.7-code-highspeed`, `kimi-k2.6`, `kimi-k2.5` models all provide a 256K context window. ### Long-Thinking Capabilities * Kimi K2.7 Code still has strong reasoning capabilities, supporting multi-step tool invocation and reasoning, excelling at solving complex problems, such as complex logical reasoning, mathematical problems, and code writing. * Kimi K2.7 Code does not support non-thinking mode. ## Example Usage Here is a complete usage example to help you quickly get started with the Kimi K2.7 Code model. ### Install the OpenAI SDK Kimi API is fully compatible with OpenAI's API format. You can install the OpenAI SDK as follows: ```bash theme={null} pip install --upgrade 'openai>=1.0' ``` ### Verify the Installation ```bash theme={null} python -c 'import openai; print("version =",openai.__version__)' # The output may be version = 1.10.0, indicating the OpenAI SDK was installed successfully and your Python environment is using OpenAI SDK v1.10.0. ``` ## Quick Start * [Try it now](https://platform.kimi.ai/playground): Test model performance in your business scenarios through interactive operations in the Dev Workbench * [Apply for API Key](https://platform.kimi.ai/console/api-keys): Test via API call immediately ### Multimodal Tool Capability Example Kimi K2.7 Code model combines multiple capabilities. The following example demonstrates K2.7 Code's visual understanding + tool calling capabilities. First, download this sample video to your local machine, such as `~/Download/test_video.mp4` Then run the following code: ```python theme={null} import base64 import json import os import subprocess import tempfile from pathlib import Path from openai import OpenAI tools = [{ "type": "function", "function": { "name": "watch_video_clip", "description": "Watch a video file or a sub-clip of it. If start_time and end_time are not provided, the entire video will be returned.", "parameters": { "type": "object", "properties": { "path": { "type": "string", "description": "The path to the video file to watch" }, "start_time": { "type": "number", "description": "The start time of the clip in seconds (optional, defaults to 0)" }, "end_time": { "type": "number", "description": "The end time of the clip in seconds (optional, defaults to end of video)" } }, "required": ["path"] } } }] def watch_video_clip(path: str, start_time: float | None = None, end_time: float | None = None) -> list[dict]: """ Watch a video file or a sub-clip of it. Args: path: The path to the video file to watch start_time: The start time in seconds (optional, defaults to 0) end_time: The end time in seconds (optional, defaults to end of video) Returns: A list of content blocks in MultiModal Tool API format """ video_path = Path(path) if not video_path.exists(): raise FileNotFoundError(f"Video file not found: {path}") # Get video duration if needed if start_time is None and end_time is None: # Return entire video with open(path, "rb") as f: video_base64 = base64.b64encode(f.read()).decode("utf-8") return [ {"type": "video_url", "video_url": {"url": f"data:video/mp4;base64,{video_base64}"}}, {"type": "text", "text": f"Full video: {video_path.name}"} ] # Get video duration for defaults probe = subprocess.run( ["ffprobe", "-v", "quiet", "-print_format", "json", "-show_format", path], capture_output=True, text=True ) duration = float(json.loads(probe.stdout)["format"]["duration"]) start_time = start_time or 0 end_time = end_time or duration clip_duration = end_time - start_time # Extract clip with tempfile.NamedTemporaryFile(suffix=".mp4", delete=False) as tmp: tmp_path = tmp.name try: subprocess.run([ "ffmpeg", "-y", "-ss", str(start_time), "-i", path, "-t", str(clip_duration), "-c:v", "libx264", "-c:a", "aac", "-preset", "fast", "-crf", "23", "-movflags", "+faststart", "-loglevel", "error", tmp_path ], check=True) with open(tmp_path, "rb") as f: video_base64 = base64.b64encode(f.read()).decode("utf-8") return [ {"type": "video_url", "video_url": {"url": f"data:video/mp4;base64,{video_base64}"}}, {"type": "text", "text": f"Clip from {video_path.name}: {start_time}s - {end_time}s"} ] finally: if os.path.exists(tmp_path): os.unlink(tmp_path) client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1" ) def agent_loop(user_message: str): """Simple agent loop with multimodal tool support.""" messages = [ {"role": "system", "content": "You are a video analysis assistant. Use watch_video_clip to examine specific portions of videos."}, {"role": "user", "content": user_message} ] while True: response = client.chat.completions.create( model="kimi-k2.7-code", messages=messages, tools=tools, tool_choice="auto" ) message = response.choices[0].message messages.append(message.model_dump()) # No tool calls = done if not message.tool_calls: return message.content # Execute tool calls for tool_call in message.tool_calls: if tool_call.function.name == "watch_video_clip": args = json.loads(tool_call.function.arguments) result = watch_video_clip( path=args["path"], start_time=args.get("start_time"), end_time=args.get("end_time") ) # Multimodal tool result messages.append({ "role": "tool", "tool_call_id": tool_call.id, "content": result }) # Usage answer = agent_loop("Analyze what happens between seconds 8-13 in ~/Download/test_video.mp4") print(answer) ``` ## Best Practices ### Supported Formats Images are supported in formats: png, jpeg, webp, gif.
Videos are supported in formats: mp4, mpeg, mov, avi, x-flv, mpg, webm, wmv, 3gpp. ### Token Calculation and Billing Image and video token usage is dynamically calculated. You can use the [token estimation API](/docs/api/estimate) to check the expected token consumption for a request containing images or video before processing. Generally, the higher the resolution of an image, the more tokens it will consume. For videos, the number of tokens depends on the number of keyframes and their resolution—the more keyframes and the higher their resolution, the greater the token consumption. The Vision model uses the same billing method as the `moonshot-v1` model series, with charges based on the total number of tokens processed. For more information, see: For token pricing details, refer to [Model Pricing](/docs/pricing/chat-k27-code). ### Recommended Resolution We recommend that image resolution should not exceed 4k (4096×2160), and video resolution should not exceed FHD (1920×1080). Higher resolutions will only increase processing time and will not improve the model's understanding. ### Upload File or Base64? Due to the limitation on the overall size of the request body, for very large videos you **must** use the file upload method to utilize vision capabilities.For images or videos that will be referenced multiple times, it is recommended to use the file upload method. Regarding file upload limitations, please refer to the [File Upload documentation](/docs/api/files-upload). Image quantity limit: The Vision model has no limit on the number of images, but ensure that the request body size does not exceed 100M URL-formatted images: Not supported, currently only supports base64-encoded image content ## Parameters Differences in Request Body Parameters are listed in [chat](/docs/api/chat). However, behaviour of some parameters may be different in k2.7-code/k2.6/k2.5 models. **We recommend using the default values instead of manually configuring these parameters.** Differences are listed below. | Field | Required | Description | Type | Values | | ------------------ | -------- | ---------------------------------------------------------------------------- | ------ | --------------------------------------------------------------------------------------------------------------- | | max\_tokens | optional | The maximum number of tokens to generate for the chat completion. | int | Default to be 32k aka 32768 | | thinking | optional | **New!** This parameter controls if the thinking is enabled for this request | object | Default to be `{"type": "enabled"}`. Kimi K2.7 Code model will throw an error if the thinking mode is disabled. | | temperature | optional | The sampling temperature to use | float | Kimi K2.7 Code model will use a fixed value 1.0. Any other value will result in an error | | top\_p | optional | A sampling method | float | Kimi K2.7 Code model will use a fixed value 0.95. Any other value will result in an error | | n | optional | The number of results to generate for each input message | int | Kimi K2.7 Code model will use a fixed value 1. Any other value will result in an error | | presence\_penalty | optional | Penalizing new tokens based on whether they appear in the text | float | Kimi K2.7 Code model will use a fixed value 0.0. Any other value will result in an error | | frequency\_penalty | optional | Penalizing new tokens based on their existing frequency in the text | float | Kimi K2.7 Code model will use a fixed value 0.0. Any other value will result in an error | ## Tool Use Compatibility When using tools, please note the following constraints to ensure model performance: * `tool_choice` can only be set to "auto" or "none" (default is "auto") to avoid conflicts between reasoning content and the specified tool\_choice. Any other value will result in an error; * During multi-step tool calling, you must keep the `reasoning_content` from the assistant message in the current turn's tool call within the context, otherwise an error will be thrown; ## Model Pricing For token pricing details, refer to [Model Pricing](/docs/pricing/chat-k27-code). ## Learn More * For the benchmark testing with Kimi K2.7 Code, please refer to this [benchmark best practice](/docs/guide/benchmark-best-practice) * See how to use Kimi K2 Series models in [Claude Code](/docs/guide/claude-code-kimi) * Learn how to configure and use the [Thinking Model](/docs/guide/use-thinking-models) * For all model pricing see [here](/docs/pricing/chat), [Billing & Rate Limit details](/docs/pricing/limits), and [Web Search Pricing](/docs/pricing/tools) # Kimi K3 Source: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart Explore Kimi K3 for long-horizon coding, knowledge work, deep reasoning, visual understanding, and a 1M-token context window. ## Introducing Kimi K3 Kimi K3 is Kimi's most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals, with native visual understanding and a 1M-token context window. It is the world's first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning. For complete benchmarks and case studies, see the [technical blog](https://www.kimi.com/blog/kimi-k3). Kimi is currently working closely with inference partners and open-source maintainers to align technical details and ensure the model launches reliably across the ecosystem. The full model weights will be released by July 27, 2026. More details on architecture, training, and evaluation will be published with the Kimi K3 technical report. ### A 3-trillion-scale open-source model Kimi K3 is the first open-source model to reach 2.8 trillion parameters. This is the latest step in Kimi's continued push of model-scale boundaries: in 9 of the past 12 months (2025/07–2026/07), Kimi models have maintained the frontier in open-source model scale. Open-source frontier model scale over time Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). Both architectural updates are designed to help information flow more smoothly through longer sequences and deeper models. We also further increased the sparsity of the Mixture of Experts (MoE): with the Stable LatentMoE framework, the model efficiently activates 16 out of 896 experts. Together with improvements in training methodology and data recipes, these structural advances give Kimi K3 roughly 2.5x the overall scaling efficiency of K2, converting compute into capability more effectively. Kimi K3 architecture ### Coding Kimi K3 has strong long-horizon coding capabilities. With minimal human supervision, it can sustain long-running engineering tasks, understand and work with large codebases, and coordinate terminal tools. Kimi K3 also excels at tasks that combine software engineering and visual reasoning. It can use screenshots and visual feedback to improve workflows in game development, frontend engineering, CAD, and related scenarios. ### Knowledge work Kimi K3 advances end-to-end knowledge work. Beyond public benchmarks, Kimi K3 (max) also shows consistent gains in our internal evaluations. These evaluations reflect recurring task patterns and challenges from real user-agent collaboration workflows. Kimi K3 demonstrates consistent advantages across production-oriented workflows, indicating broad improvements in agentic knowledge-work capabilities. ## Access requirements Kimi K3 is a flagship model: it is unlocked after a successful top-up (minimum \$1). Your cumulative top-up amount also determines your account tier and rate limits (concurrency, RPM, TPM, TPD) — see [Recharge and Rate Limits](/docs/pricing/limits). ## Get started * [Playground](https://platform.kimi.ai/playground) * [Get an API Key](https://platform.kimi.ai/console/api-keys) The examples require Python 3.9+ and the OpenAI SDK. Install the SDK and initialize the client once; later Python examples reuse `client`. ```bash theme={null} python3 -m pip install --upgrade 'openai>=1.0' ``` ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) ``` ## Basic call ```python theme={null} completion = client.chat.completions.create( model="kimi-k3", messages=[{"role": "user", "content": "Introduce Kimi K3 in one sentence."}], ) print(completion.choices[0].message.content) ``` ```bash theme={null} curl https://api.moonshot.ai/v1/chat/completions \ --header "Authorization: Bearer $MOONSHOT_API_KEY" \ --header "Content-Type: application/json" \ --data '{ "model": "kimi-k3", "messages": [{"role": "user", "content": "Introduce Kimi K3 in one sentence."}] }' ``` ## Reasoning effort K3 always has thinking mode enabled and supports configuring its reasoning effort with the top-level `reasoning_effort` request field. Reasoning effort supports `low`, `high`, and `max` (default `max`). See [Reasoning Effort](/docs/guide/use-reasoning-effort) for usage. ```python theme={null} completion = client.chat.completions.create( model="kimi-k3", reasoning_effort="max", messages=[{"role": "user", "content": "Prove that the square root of 2 is irrational."}], ) print(completion.choices[0].message.content) ``` For multi-turn conversations and tool calls, add the complete assistant message returned by the API to the next request. Do not keep only `content`. ## Streaming Streaming responses provide separate `reasoning_content` and final-answer `content` deltas. See [Streaming Output](/docs/guide/utilize-the-streaming-output-feature-of-kimi-api) for details. ```python theme={null} stream = client.chat.completions.create( model="kimi-k3", messages=[{"role": "user", "content": "Explain why the sky is blue."}], stream=True, ) for chunk in stream: delta = chunk.choices[0].delta reasoning = getattr(delta, "reasoning_content", None) if reasoning: print(reasoning, end="", flush=True) if delta.content: print(delta.content, end="", flush=True) ``` ## Vision input For vision messages, `content` must be an array of objects, not a serialized string. See [Vision Input](/docs/guide/use-kimi-vision-model) for formats and limits. ```python theme={null} import base64 from pathlib import Path image_data: str = base64.b64encode(Path("image.png").read_bytes()).decode() completion = client.chat.completions.create( model="kimi-k3", messages=[ { "role": "user", "content": [ { "type": "image_url", "image_url": {"url": f"data:image/png;base64,{image_data}"}, }, {"type": "text", "text": "Describe this image."}, ], } ], ) print(completion.choices[0].message.content) ``` ```python theme={null} from pathlib import Path video = client.files.create(file=Path("video.mp4"), purpose="video") try: completion = client.chat.completions.create( model="kimi-k3", messages=[ { "role": "user", "content": [ { "type": "video_url", "video_url": {"url": f"ms://{video.id}"}, }, {"type": "text", "text": "Summarize this video."}, ], } ], ) print(completion.choices[0].message.content) finally: client.files.delete(video.id) ``` ## Structured output Use `json_schema` with `strict: true` to constrain the final `message.content`. Parse only that field, not `reasoning_content`. ```python theme={null} import json completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "user", "content": "Lin is 28 years old. Extract the name and age."} ], response_format={ "type": "json_schema", "json_schema": { "name": "person", "strict": True, "schema": { "type": "object", "properties": { "name": {"type": "string"}, "age": {"type": "integer"}, }, "required": ["name", "age"], "additionalProperties": False, }, }, }, ) person: dict[str, object] = json.loads( completion.choices[0].message.content or "{}" ) print(person) ``` See [Structured Output](/docs/guide/response_format). ## Partial Mode Add an assistant message with `partial=True` at the end of `messages` to continue from a text prefix. Prepend that prefix when displaying the final result. ```python theme={null} prefix: str = "Conclusion: " completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "user", "content": "In one sentence, explain why API compatibility matters."}, {"role": "assistant", "content": prefix, "partial": True}, ], ) print(prefix + (completion.choices[0].message.content or "")) ``` See [Partial Mode](/docs/guide/use-partial-mode-feature-of-kimi-api). ## Custom tools and `tool_choice` Use `tool_choice="required"` on the first turn to require at least one tool call. After executing every call, return the complete assistant message and append one tool result with the matching `tool_call_id` for each call. ```python theme={null} import json from typing import Any tools: list[dict[str, Any]] = [ { "type": "function", "function": { "name": "get_weather", "description": "Get the weather for a city", "parameters": { "type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"], }, }, } ] messages: list[Any] = [ {"role": "user", "content": "What is the weather in San Francisco today?"} ] first = client.chat.completions.create( model="kimi-k3", messages=messages, tools=tools, tool_choice="required", ) assistant_message = first.choices[0].message messages.append(assistant_message) for tool_call in assistant_message.tool_calls or []: arguments: dict[str, str] = json.loads(tool_call.function.arguments) result: str = json.dumps( {"city": arguments["city"], "weather": "sunny", "temperature_c": 24} ) messages.append( {"role": "tool", "tool_call_id": tool_call.id, "content": result} ) final = client.chat.completions.create( model="kimi-k3", messages=messages, tools=tools, ) print(final.choices[0].message.content) ``` See [Tool Choice](/docs/guide/use-tool-choice). ## Dynamic tool loading Place a complete tool definition in a `system` message without `content`. The tool becomes available from that message onward. ```python theme={null} from typing import Any dynamic_messages: list[dict[str, Any]] = [ {"role": "user", "content": "Calculate 23 times 47."}, { "role": "system", "tools": [ { "type": "function", "function": { "name": "calculate", "description": "Evaluate an arithmetic expression", "parameters": { "type": "object", "properties": { "expression": { "type": "string", "description": "The arithmetic expression to evaluate", } }, "required": ["expression"], }, }, } ], }, ] completion = client.chat.completions.create( model="kimi-k3", messages=dynamic_messages, ) print(completion.choices[0].message.tool_calls) ``` * Include the complete `name`, `description`, and `parameters` definition. * The declaration takes effect at its position in `messages`. * Keep this message in later request history; the server does not retain it. See [Dynamic Tool Loading](/docs/guide/use-dynamic-tool-loading). ## 1M context and automatic caching A new request can hit the prefix cache only when the previous request's prompt tokens exceed 256. If the previous request's prompt tokens are below 256, the request is not cached and is discarded. See [Context Caching](/docs/guide/use-context-caching-feature-of-kimi-api) for details. Context caching is automatic for regular model requests; no cache ID, TTL, or extra parameter is required. Keep the long prefix unchanged so later requests can automatically attempt a cache hit. ```python theme={null} from pathlib import Path knowledge: str = Path("knowledge-base.md").read_text(encoding="utf-8") for question in ["Summarize the key conclusions.", "List three implementation risks."]: completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "system", "content": knowledge}, {"role": "user", "content": question}, ], ) print(completion.choices[0].message.content) ``` See [Context Caching](/docs/guide/use-context-caching-feature-of-kimi-api). ## Official tools Official tools are integrated through Formula: 1. Fetch tool definitions from the Formula `/tools` endpoint. 2. Add those definitions to the Chat Completions `tools` field. 3. When the model returns `tool_calls`, submit each function name and arguments to the Formula `/fibers` endpoint. 4. Add the complete assistant message and Fiber output as the corresponding tool message. 5. Call Chat Completions again until the model returns a final answer. See [Official Tools](/docs/guide/use-official-tools) for the complete client and API contract. Web search is being updated and is not recommended for use in the near term. ## Important limits * Reasoning effort is configured with the top-level `reasoning_effort` request field and supports `low`, `high`, and `max` (default `max`); K3 always has thinking mode enabled. * `max_completion_tokens` defaults to 131072 and can be set up to 1048576. * `temperature=1.0`, `top_p=0.95`, `n=1`, `presence_penalty=0`, and `frequency_penalty=0` are fixed; omit them from requests. * Return the complete assistant message unchanged in multi-turn conversations and tool calls. * Vision input does not support public image URLs. Use base64 or `ms://`, and make `content` an array of objects. * Web search is being updated and is not recommended for production workflows in the near term. ## FAQ Kimi K3 offers a 1M-token context and uses flat pay-as-you-go pricing — there is no tiering by context length. Input (with separate rates for cache hits and misses) and output are billed at uniform per-token prices. See [Kimi K3 pricing](/docs/pricing/chat-k3). You can't — K3 always thinks. If the reasoning takes too long, set `reasoning_effort` to `low` to reduce the reasoning effort. See [Reasoning Effort](/docs/guide/use-reasoning-effort). ## Model Pricing For token pricing details, refer to [Model Pricing](/docs/pricing/chat-k3). Dear Kimi users: Due to a recent increase in high-frequency abnormal requests on the platform, which has affected the stability of cluster services, we plan to update the “Top-up Tiers and Rate Limits” rules in August. Please visit the [Recharge and Rate Limits](/docs/pricing/limits) page for updates. ## Related docs Configure reasoning\_effort. Send images and videos. Use strict JSON Schema. Continue from a prefix. Control whether the model calls tools. Inject tool definitions on demand. Combine tool-calling features. Integrate Formula tools. Review input and output prices. # Kimi K3 API Tool Calling Best Practices Source: https://platform.kimi.ai/docs/guide/kimi-k3-tool-calling-best-practice When your agent has a large tool inventory, combine dynamic loading, tool_choice, and reasoning effort in the tool-calling flow. When your agent has access to dozens or hundreds of tools, don't put every tool definition into the request — they eat up context and make the model more likely to pick the wrong tool. This guide walks through a tool-orchestration setup on Kimi K3: retrieve candidate tools with a search tool first, then inject tool definitions into the conversation on demand. ## Declare a search tool, not all your tools At the start of a conversation, declare only a single `search_tools` function — implemented by your backend — plus a small set of core tools you expect to use in every turn: ```json theme={null} { "tools": [ { "type": "function", "function": { "name": "search_tools", "description": "Search available tools by keyword and return matching tool names and summaries", "parameters": { "type": "object", "properties": { "query": { "type": "string", "description": "Search keyword, e.g. github or database" } }, "required": ["query"] } } } ] } ``` In the system prompt, advertise the domain tags the model can search (for example, a tool catalog or business domains) so it knows to call `search_tools` first when it needs a tool. No matter how large your total inventory is, each request then only carries a handful of tool declarations. ## Use tool\_choice to force first-turn retrieval The model may choose not to call any tool and answer from memory. To make sure it retrieves before answering, set `tool_choice: "required"` on the first turn: ```json theme={null} { "model": "kimi-k3", "messages": [{"role": "user", "content": "Help me create a GitHub PR"}], "tools": ["..."], "tool_choice": "required" } ``` After retrieval, switch `tool_choice` back to `"auto"` for subsequent requests. Changing `tool_choice` does not invalidate the prefix cache, so you can adjust it per request. See [Tool Choice](/docs/guide/use-tool-choice) for all accepted values. ## Inject tool definitions on demand When `search_tools` returns candidate tools, your application inserts the full declarations of the matching tools into `messages` via a `system` message carrying a `tools` field. The tools become visible to the model starting from that message's position: ```json theme={null} { "role": "system", "tools": [ { "type": "function", "function": { "name": "create_github_pr", "description": "Create a pull request in the given repository", "parameters": { "type": "object", "properties": {} } } } ] } ``` Dynamic declarations use exactly the same format as the top-level `tools` field — no second schema to maintain — and coexist with globally declared tools. Dynamic tool declarations apply per request and are not retained by the server. In the next request, the client can keep the original declaration so the tool remains available and the prefix cache can be reused, or remove it. If the tool is not declared elsewhere, the model cannot call that tool, and the changed prefix may miss the cache. See [Dynamically Loaded Tools](/docs/guide/use-dynamic-tool-loading) for full usage. A new request can hit the prefix cache only when the previous request's prompt tokens exceed 256. If the previous request's prompt tokens are below 256, the request is not cached and is discarded. See [Context Caching](/docs/guide/use-context-caching-feature-of-kimi-api) for details. ## Pick the reasoning effort for the task The top-level `reasoning_effort` request field supports `low`, `high`, and `max`, with `max` as the default. Decide on this setting before the conversation starts. Appending a dynamic tool declaration to the end of `messages` does not affect the cached prefix; removing or modifying an earlier tool declaration may affect cache hits after the point of change. Changing `tool_choice` does not invalidate the prefix cache. See [Reasoning Effort](/docs/guide/use-reasoning-effort) for configuration details. ## The complete flow 1. Conversation start: top-level `tools` carries only `search_tools` plus a few core tools; 2. First-turn retrieval: `tool_choice: "required"` forces the model to call `search_tools`; 3. Inject on demand: insert tool definitions via a `system` message based on the retrieval results; 4. Call directly: the model calls the loaded tools in subsequent generations; 5. Reasoning effort: decide on the top-level `reasoning_effort` setting before the conversation starts. ## Related reading * [Dynamically Loaded Tools](/docs/guide/use-dynamic-tool-loading) * [Tool Choice](/docs/guide/use-tool-choice) * [Reasoning Effort](/docs/guide/use-reasoning-effort) * [Use Kimi API for Tool Calls](/docs/guide/use-kimi-api-to-complete-tool-calls) * [Model Parameter Reference](/docs/api/models-overview) # Use Kimi Models in OpenCode Source: https://platform.kimi.ai/docs/guide/open-code Install OpenCode, connect it to the Kimi Open Platform through built-in authentication, and use Kimi K3 with its thinking-effort variants. [OpenCode](https://opencode.ai/) is an open-source programming agent. This guide shows how to connect OpenCode to the Kimi Open Platform through built-in authentication and use the `kimi-k3` model with its 1M-token context window. This guide is based on OpenCode 1.18.3. Its interface, configuration options, and supported capabilities may change between versions. ## Prerequisites Complete these prerequisites before you start. Follow the linked official guides for installation and account setup. Install or update OpenCode with the official documentation. Create a Kimi Open Platform API key and keep it private. Confirm that the account has available balance, then check limits, budgets, and organization settings. Kimi K3 requires available balance in your account; vouchers granted from new-user verification cannot be used for Kimi K3. Rate limits vary by user tier. See [Rate Limits](/docs/pricing/limits). If your organization uses an IP whitelist, follow [Organization Best Practices](/docs/guide/org-best-practice) and add the public egress IPv4 address of the computer that calls the API. ## Step 1: Configure the API key Run `opencode auth login` and select **Moonshot AI** in the provider list: ```text theme={null} $ opencode auth login ┌ Add credential │ ◆ Select provider │ Search: Moon█ (2 matches) │ ● Moonshot AI │ ↑/↓ to select • Enter: confirm • Type: to search └ ``` Then paste your Kimi Open Platform API key and press `Enter`: ```text theme={null} $ opencode auth login ┌ Add credential │ ◇ Select provider │ Moonshot AI │ ◇ Enter your API key │ ▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪▪ │ └ Done ``` Do not put the API key in configuration files, screenshots, or a Git repository. This guide uses a Global Kimi Open Platform key; do not substitute a key from another Kimi service or region. ## Step 2: Select the Kimi K3 model Run `opencode` to start OpenCode: ```bash theme={null} opencode ``` Run the `/models` command in the input field: Run the /models command Search for and select **Kimi K3** in the Select model dialog: Select the Kimi K3 model ## Step 3: Adjust the reasoning effort Run the `/variants` command in the input field: Run the /variants command Select **max** in the Select variant dialog: Select the max variant `kimi-k3` defaults to `max` reasoning effort and can also switch to `low` / `high`. See the [Model Parameter Reference](/docs/api/models-overview). When setup is complete, the status bar should show **Kimi K3**, **Moonshot AI**, and **max**: Final configuration ## Learn more * [OpenCode documentation](https://opencode.ai/docs) * [Kimi Open Platform quickstart](/docs/overview) # Setting Up and Verifying Your Organization Source: https://platform.kimi.ai/docs/guide/org-best-practice Create and verify a Kimi Open Platform organization, configure an IP allowlist, and manage members, projects, and API keys. When you register and log in to the Open Platform account, you can find your organization ID on the **Organization Management** - **Organization Verification** page. The organization ID is the unique identifier for your organization. ## **Configure IP Whitelist** IP Whitelist is an organization-level security setting on Kimi Open Platform. After organization verification is completed, you can configure an IP whitelist. Once saved, only IP addresses in the whitelist can access APIs under the current organization. API requests from IP addresses outside the whitelist will be denied. When the IP whitelist is empty, API requests are not restricted by source IP. When saved, the current list overwrites the existing configuration. ### **Entry and Permissions** After logging in to [Kimi Open Platform](https://platform.kimi.ai/), go to **Organization Management** - **Organization Verification** from the left sidebar. After verification is completed and the current account has permission, the **IP Whitelist** entry will be displayed in the upper-right corner of the page. Image The entry and configuration permissions are as follows: | **Organization Status** | **Entry Displayed** | **Who Can Configure** | | :-------------------------------------- | :------------------ | :----------------------------------------------- | | Organization verification not completed | No | None | | Organization verification completed | Yes | Current organization account | | Enterprise Verified | Yes | Organization creator, organization administrator | If you do not see the **IP Whitelist** entry, first check whether organization verification has been completed. For enterprise organizations, only the organization creator and organization administrators can see and configure the entry. ### **Fill in the IP / CIDR List** Click **IP Whitelist**, then fill in the allowed IP addresses or CIDR ranges in the **IP / CIDR List** field. 图片 Please note: * Enter one item per line. You can also separate items with commas or spaces. * You can configure up to 20 items. * Only public IPv4 addresses and valid IPv4 CIDR ranges are supported. * IPv6 is not currently supported. * Saving the whitelist overwrites the existing configuration. Example: ```text theme={null} 203.0.113.4 198.51.100.0/24 52.94.76.0/22 13.107.42.0/24 ``` You can also write them as: ```text theme={null} 203.0.113.4, 198.51.100.0/24 52.94.76.0/22 13.107.42.0/24 ``` Before configuring the whitelist, make sure your service uses a fixed public IPv4 egress address. If your service is deployed behind a cloud provider, proxy gateway, NAT gateway, or corporate network, enter the actual public egress IP address used when calling the Kimi API. ### **Format Validation** The system validates the input before saving. The following items cannot be saved: * IPv6 addresses; * private IPv4 addresses or private IPv4 ranges; * loopback addresses, link-local addresses, multicast addresses, reserved addresses, and other non-public IPv4 addresses; * invalid CIDR ranges; * lists with more than 20 items. Common invalid examples: | **Type** | **Examples** | | :------------------------------- | :----------------------------------------- | | Private address or range | `192.168.1.1`, `10.0.0.1`, `172.16.0.0/12` | | Carrier-grade NAT shared address | `100.64.0.0/10` | | Loopback address | `127.0.0.0/8` | | Link-local address | `169.254.0.0/16` | | Multicast or reserved address | `224.0.0.0/4`, `240.0.0.0/4` | If the input contains invalid items, the dialog will show an error message and mark the invalid entries. Fix the issues before saving. ### **Clear the IP Whitelist** To remove IP whitelist restrictions, open the **IP Whitelist** dialog, clear all content in the **IP / CIDR List** field, and click **Save**. After the empty list is saved, API access under the current organization will no longer be restricted by IP whitelist checks. ### **Scope** The IP whitelist applies to API access under the current organization. After it is configured, API Keys under the current organization are subject to this whitelist when calling APIs. If your business has multiple network egress points, add all required public egress IP addresses or CIDR ranges to the whitelist. ## Organization Balance Alert To prevent service interruptions due to insufficient account balance, we recommend configuring balance alerts in the [Organization Management Settings](https://platform.kimi.ai/console/account). * The platform provides customizable balance alert thresholds, with a default setting of \$5 * When your account balance drops below the configured threshold, the system will automatically send notification emails to the organization's registered email address settings ## Managing Projects and Usage Limits To meet the needs of multiple business product lines under a single organization, or to distinguish between production and testing environments, you can create multiple projects under your organization. Within each project, you can create an API Key. The calls made using the project's API Key will be recorded under the project's consumption, allowing you to independently manage the usage of different projects. ### Project Balance and Rate Limiting * All projects under an organization share the organization's rate limits. * All projects under an organization share the organization's account balance. ### Project Consumption Management * The platform now supports setting monthly and daily consumption budgets on a per-project basis. You can set the monthly or daily consumption limits for each project on the **Project Management** - **Project Settings** - **Project Budget/Rate Limiting Settings** page. Once the API Key consumption within a project reaches the set budget, any subsequent API requests for that project will be denied, effectively helping you manage your business budget. Due to billing cycle issues, the actual enforcement of these limits may have a delay of about 10 minutes. settings * If you wish to limit the maximum TPM (Transactions Per Minute) for a single project, you can configure the project's TPM rate limit independently. If the project's API Key requests reach this TPM, the requests will be denied. (The project's TPM must not exceed the organization's TPM. If you set a value higher than the organization's TPM, the organization's TPM will be used for rate limiting.) * The platform also provides an overview page for both the organization and individual projects, offering consumption analysis at both levels to help you get a clear understanding of your organization's spending. ### Project Quantity Limitations The number of projects your organization can create depends on the type of organization verification. The upper limits for different verification types are as follows: | Organization Type | Project Limit | API Key Limit | | ------------------- | ------------- | ------------- | | Default | 20 | 50 | | Enterprise Verified | 50 | 100 | If you have additional requirements, please fill out the [Contact Sales](https://platform.kimi.ai/console/contact-sales) form for consultation. ## Member Management ### Organization Member Management To help you manage your organization, you can invite new members on the **Organization Management** - **Member Management** page. The platform generates a dedicated invitation link for each new member. The invitee can use this link to register and log in to the Open Platform and join your organization.\ **Note:** Please visit the [Set Organization Information](https://platform.kimi.ai/console/general) page to maintain your organization's info and complete your enterprise verification as a prerequisite. invite invite1 * **Organization Administrator**: The organization creator is the default administrator. Administrators can create projects, invite and manage members, and issue invoices. * **General Member**: General members can only view projects. They must be invited to join a project to gain access to project resources. ### Project Member Management Organization administrators can create projects and invite organization members to join and help manage projects. Project members can create their own API Keys within the project to utilize project resources. * **Project Administrator**: Can manage project budget/rate limits/consumption notifications, invite members, and create API Keys. * **General Project Member**: Can only view projects / create API Keys. ### Project API Key Management It is recommended that each project member creates their own API Key within a project rather than sharing keys. When a member is removed from a project, all API Keys created by that member will also be invalidated, helping your organization effectively manage project resources. # Compare with Other Kimi Products Source: https://platform.kimi.ai/docs/guide/product-plans The Kimi API Open Platform uses pay-as-you-go billing with no subscription plan. It is different from products such as Kimi Membership and Kimi Code — please distinguish between them. # Best Practices for Prompts Source: https://platform.kimi.ai/docs/guide/prompt-best-practice Write more reliable and controllable Kimi system and user prompts with clear instructions, examples, roles, and output constraints. > Best Practices for System Prompts: A system prompt refers to the initial input or instruction that a model receives before generating text or responding. This prompt is crucial for the model's operation [link](https://kimi.moonshot.cn/share/col3fn2lnl95v16j0g2g). ## Write Clear Instructions * Why is it necessary to provide clear instructions to the model? > The model can't read your mind. If the output is too long, you can ask the model to respond briefly. If the output is too simple, you can request expert-level writing. If you don't like the format of the output, show the model the format you'd like to see. The less the model has to guess about your needs, the more likely you are to get satisfactory results. ### Including More Details in Your Request Can Yield More Relevant Responses > To obtain highly relevant output, ensure that your input request includes all important details and context. | General Request | Better Request | | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | How to add numbers in Excel? | How do I sum a row of numbers in an Excel table? I want to automatically sum each row in the entire table and place all the totals in the rightmost column named "Total." | | Work report summary | Summarize my work records from 2023 in a paragraph of no more than 500 words. List the highlights of each month in sequence and provide a summary of the entire year. | ### Requesting the Model to Assume a Role Can Yield More Accurate Output > Add a specified role for the model to use in its response in the 'messages' field of the API request. ```json theme={null} { "messages": [ {"role": "system", "content": "You are Kimi, an artificial intelligence assistant provided by Moonshot AI. You are more proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. At the same time, you will refuse to answer any questions involving terrorism, racism, or explicit violence. Moonshot AI is a proper noun and should not be translated into other languages."}, {"role": "user", "content": "Hello, my name is Li Lei. What is 1+1?"} ] } ``` ### Using Delimiters in Your Request to Clearly Distinguish Different Parts of the Input > For example, using triple quotes/XML tags/section headings as delimiters can help distinguish text parts that require different processing. ```json theme={null} { "messages": [ {"role": "system", "content": "You will receive two articles of the same category, separated by XML tags. First, summarize the arguments of each article, then point out which article presents a better argument and explain why."}, {"role": "user", "content": "
Insert article here
Insert article here
"} ] } ``` ```json theme={null} { "messages": [ {"role": "system", "content": "You will receive an abstract and the title of a paper. The title should give readers a clear idea of the paper's topic and also be eye-catching. If the title you receive does not meet these standards, please suggest five alternative options."}, {"role": "user", "content": "Abstract: Insert abstract here.\n\nTitle: Insert title here"} ] } ``` ### Clearly Define the Steps Needed to Complete the Task > It is advisable to outline a series of steps for the task. Writing these steps explicitly makes it easier for the model to follow and produces better output. ```json theme={null} { "messages": [ {"role": "system", "content": "Respond to user input using the following steps.\nStep one: The user will provide text enclosed in triple quotes. Summarize this text into one sentence with the prefix “Summary: ”.\nStep two: Translate the summary from step one into English and add the prefix \"Translation: \"."}, {"role": "user", "content": "\"\"\"Insert text here\"\"\""} ] } ``` ### Provide Examples of Desired Output to the Model > Providing examples of general guidance is usually more efficient for the model's output than showing all permutations of the task. For instance, if you intend to have the model replicate a style that is difficult to describe explicitly in response to user queries, this is known as a "few-shot" prompt. ```json theme={null} { "messages": [ {"role": "system", "content": "Respond in a consistent style"}, {"role": "user", "content": "Insert text here"} ] } ``` ### Specify the Desired Length of the Model's Output > You can request the model to generate output of a specific target length. The target output length can be specified in terms of words, sentences, paragraphs, bullet points, etc. However, note that instructing the model to generate a specific number of words is not highly precise. The model is better at generating output of a specific number of paragraphs or bullet points. ```json theme={null} { "messages": [ {"role": "user", "content": "Summarize the text within the triple quotes in two sentences, within 50 words. \"\"\"Insert text here\"\"\""} ] } ``` ## Provide Reference Text ### Guide the Model to Use Reference Text to Answer Questions > If you can provide a model with credible information related to the current query, you can guide the model to use the provided information to answer the question. ```json theme={null} { "messages": [ {"role": "system", "content": "Answer the question using the provided article (enclosed in triple quotes). If the answer is not found in the article, write \"I can't find the answer.\""}, {"role": "user", "content": ""} ] } ``` ## Break Down Complex Tasks ### Categorize to Identify Instructions Relevant to User Queries > For tasks that require a large set of independent instructions to handle different scenarios, categorizing the query type and using this categorization to clarify which instructions are needed may aid the output. Based on the classification of the customer query, a set of more specific instructions can be provided to the model to help it handle subsequent steps. For example, assume the customer needs help with “troubleshooting.” ```json theme={null} { "messages": [ {"role": "system", "content": "You will receive a customer service inquiry that requires technical support. You can assist the user in the following ways:\n\n-Ask them to check if *** is configured.\nIf all *** are configured but the problem persists, ask for the device model they are using\n-Now you need to tell them how to restart the device:\n=If the device model is A, perform ***.\n-If the device model is B, suggest they perform ***."} ] } ``` ### For Long-Running Dialog Applications, Summarize or Filter Previous Conversations > Since the model has a fixed context length, the conversation between the user and the model assistant cannot continue indefinitely. One solution to this issue is to summarize the first few rounds of the conversation. Once the input size reaches a predetermined threshold, a query is triggered to summarize the previous part of the conversation, and the summary of the previous conversation can also be included as part of the system message. Alternatively, previous conversations throughout the entire chat process can be summarized asynchronously. ### Chunk and Recursively Build a Complete Summary for Long Documents > To summarize the content of a book, we can use a series of queries to summarize each chapter of the document. Partial summaries can be aggregated and summarized to produce a summary of summaries. This process can be recursively repeated until the entire book is summarized. If understanding later parts requires reference to earlier chapters, then when summarizing a specific point in the book, include summaries of the chapters preceding that point. # Use response_format to control model output format Source: https://platform.kimi.ai/docs/guide/response_format Use `response_format` for JSON Mode or Structured Output and handle schema, structure, and Partial Mode constraints. The Kimi API constrains the output format of chat completions via the `response_format` parameter. It supports two modes: | Mode | `type` value | Description | Use case | | --------------------- | ------------- | -------------------------------------------------------------------------- | ---------------------------------------------------------------------- | | **JSON Mode** | `json_object` | Guarantees a valid JSON Object, but does not constrain specific fields | Simple JSON output, flexible-field scenarios | | **Structured Output** | `json_schema` | Precisely defines field names, types, and nested structure via JSON Schema | Scenarios requiring strict structure and downstream-system integration | This document focuses on the **`json_schema` mode of `response_format` (i.e., Structured Output)**, including parameter usage, model differences, common issues, and error handling. For the basics of JSON Mode, see [JSON Mode](/docs/guide/use-json-mode-feature-of-kimi-api). ## response\_format basic structure ```python theme={null} response_format={ "type": "json_schema", # or "json_object" "json_schema": { # required for json_schema mode "name": "schema_name", "strict": True, "schema": { ... } # your JSON Schema } } ``` * When `type` is `json_object`, the `json_schema` field is not required. * When `type` is `json_schema`, both `json_schema.name` and `json_schema.schema` are required. ## Advantages of Structured Output Compared to JSON Mode, Structured Output offers the following advantages: * **Strictly controlled structure**: The model output must fully follow the JSON Schema you define, with field names, types, and nesting levels matching one-to-one. * **No need to repeatedly describe the format in the prompt**: Decouple format requirements from the schema, reducing prompt-engineering complexity. * **More reliable downstream integration**: Output can be directly parsed by `json.loads` into strongly-typed objects without extra fault-tolerance handling. > **Model-difference note**: Different models have different levels of JSON Schema support. > > * `kimi-k3` reliably supports Structured Output, including nested objects, arrays, and `anyOf`. > * `kimi-k2.7-code` has the most stable Structured Output support, including nested objects, arrays, `anyOf` / `oneOf` / `$ref` / `additionalProperties: true`, etc. > * `kimi-k2.6` occasionally behaves unstably with complex schemas; for example, `$ref` may return a Markdown code block, `oneOf` may be ignored, and `partial=true` may output fields outside the schema. When using `kimi-k2.6`, prefer simple schemas and add a second validation layer in your business logic. ## Quick start ### Basic usage Set `type` to `"json_schema"` in `response_format`, and pass the `json_schema` object: ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[ { "role": "system", "content": "You are a news summarization assistant." }, { "role": "user", "content": "Please summarize the following news: Today, the field of artificial intelligence technology has welcomed a major breakthrough..." } ], response_format={ "type": "json_schema", "json_schema": { "name": "news_summary", "strict": True, "schema": { "type": "object", "properties": { "title": {"type": "string", "description": "News headline"}, "author": {"type": "string", "description": "Author or source"}, "publish_time": {"type": "string", "description": "Publication time in ISO 8601 format"}, "summary": {"type": "string", "description": "Summary within 200 characters"}, "keywords": { "type": "array", "items": {"type": "string"}, "description": "3-5 keywords" } }, "required": ["title", "author", "summary", "keywords"] } } } ) import json result = json.loads(completion.choices[0].message.content) print(result["title"]) print(result["keywords"]) ``` ### Example output ```json theme={null} { "title": "Major Breakthrough in Artificial Intelligence Technology", "author": "Tech Daily", "publish_time": "2024-06-19", "summary": "Researchers have made new progress in deep learning model efficiency optimization...", "keywords": ["artificial intelligence", "deep learning", "model optimization", "breakthrough"] } ``` ### About reasoning\_content Thinking models such as `kimi-k3` and `kimi-k2.7-code` may return `reasoning_content` in addition to `content`. Only parse `choices[0].message.content` as the final JSON; do not call `json.loads` on the entire response object. ```python theme={null} content = completion.choices[0].message.content result = json.loads(content) ``` ## Parameter description | Parameter | Type | Description | | -------------------- | ---------------------------------- | ------------------------------------------------------------------------------- | | `type` | `"json_schema"` \| `"json_object"` | Must be set; choose one | | `json_schema.name` | string | Identifier name for the schema, used for logging and debugging | | `json_schema.strict` | boolean | Whether to strictly enforce the schema. Recommended to explicitly set to `true` | | `json_schema.schema` | object | JSON Schema object defining the output structure | > **Note**: Whether `strict` is `true`, `false`, or omitted, `kimi-k2.7-code` generally adheres to the schema well; `kimi-k2.6` is more likely to output fields outside the schema when `strict=false` or omitted. Always explicitly set `strict: true`. ## `strict` mode `json_schema.strict` is recommended to be set to `true`, meaning the model output **must** fully match the schema definition. In this case, your schema must comply with the **MFJS (Moonshot Flavored JSON Schema)** specification. > **MFJS model differences**: > > * `kimi-k2.7-code` already has relatively complete support for features such as `anyOf` / `oneOf` / `$ref` / `additionalProperties: true`, and usually will not trigger MFJS errors. > * `kimi-k2.6` is more likely to hit MFJS limits with complex schemas; keep schemas simple. If `strict` is set to `false`, the API only guarantees that the output is a valid JSON object, but does not strictly enforce the internal field structure. This can be used when the schema is complex or you want to give the model more flexibility. ### How to validate schema compliance with MFJS You can use the `walle` CLI tool to quickly self-check schema compatibility: ```bash theme={null} # Install the walle tool go install github.com/moonshotai/walle/cmd/walle@latest # Validate your schema walle -schema 'your_schema_json' -level strict ``` > Even if the schema contains `anyOf` / `oneOf` / `$ref`, the API often returns `200`, and the response **does not contain a `warning` field**. Therefore, `walle` is better suited as a static-check entry point; actual compatibility should be verified via live calls against the target model. ## Nested objects and arrays example Structured Output supports arbitrarily deep nested objects and arrays, which is stable on `kimi-k2.7-code`: ```python theme={null} response_format={ "type": "json_schema", "json_schema": { "name": "meeting_minutes", "strict": True, "schema": { "type": "object", "properties": { "meeting_title": {"type": "string"}, "date": {"type": "string"}, "attendees": { "type": "array", "items": { "type": "object", "properties": { "name": {"type": "string"}, "role": {"type": "string"}, "present": {"type": "boolean"} }, "required": ["name", "role", "present"] } }, "agenda_items": { "type": "array", "items": { "type": "object", "properties": { "topic": {"type": "string"}, "discussion": {"type": "string"}, "action_items": { "type": "array", "items": { "type": "object", "properties": { "assignee": {"type": "string"}, "task": {"type": "string"}, "deadline": {"type": "string"} }, "required": ["assignee", "task"] } } }, "required": ["topic", "discussion"] } } }, "required": ["meeting_title", "date", "attendees", "agenda_items"] } } } ``` ## Comparison with JSON Mode | Feature | `json_object` | `json_schema` + `strict: true` | | ----------------- | ---------------------------------------- | -------------------------------------------------------------------------------- | | Output validity | Guaranteed valid JSON Object | Guaranteed valid JSON Object | | Field names | Not guaranteed — the model may improvise | Strictly fixed | | Field type | Not enforced | Enforced to match | | Extra fields | May be added at will | Forbidden (`additionalProperties: false`) | | Missing fields | May be omitted | `required` fields always appear; declare union types to allow `null` | | Mechanism | Prompt guidance | Token-level constrained decoding (CFG), filtering illegal tokens during sampling | | Use case | Rapid prototyping, non-critical paths | Production, API integration, data ingestion | | strict validation | None | Yes (MFJS specification) | The structural guarantee of constrained decoding holds when the schema complies with the MFJS specification; complex schemas may still be unstable on models such as `kimi-k2.6` — see the model-difference note above. For structured data consumed downstream, always use `json_schema` with `strict: true` to avoid writing extensive defensive code in the business layer. ## Notes 1. **Schema must comply with MFJS**: When `strict=true`, use the `walle` CLI tool to pre-validate the schema. Common MFJS constraints are greatly relaxed on `kimi-k2.7-code`, but may still trigger on `kimi-k2.6`. 2. **Prompt still needs context**: Although the format is constrained by the schema, the model still needs to understand the **business content**. Please clearly describe the task objective and data source in the system prompt or user prompt. 3. **`additionalProperties`**: * When set to `false`, the model will not output fields not defined in the schema. * When set to `true` or omitted, `kimi-k2.7-code` allows extra fields; `kimi-k2.6` may also output extra fields, but with less stability than `kimi-k2.7-code`. 4. **Use nullable union types for missing information**: Fields declared in `required` always appear in the output. When the input lacks the corresponding information, a field declared with a single type (e.g. `"integer"`) may lead the model to fabricate content or return an empty string. Prefer a union type that allows `null` (e.g. `"type": ["integer", "null"]`), so the model can use `null` to explicitly mean "information missing" instead of the string `"unknown"` or a disappearing field — downstream code can then `json.loads` and cast to typed objects without defensive handling. Note that `kimi-k2.6` may still return empty strings (e.g. `"employee_id": ""`), so keep a null-value check in the business layer. 5. **Error handling**: When the schema is too complex or the prompt contradicts the schema, the model may output incomplete JSON (`finish_reason="length"`). We recommend checking `finish_reason` and appropriately increasing `max_tokens`. 6. **Compatibility with Partial Mode**: * `kimi-k2.7-code` usually works with `partial=true` on simple schemas, but complex schemas may still break structural constraints. * `kimi-k2.6` is more likely to output fields outside the schema with `partial=true`, so **it is not recommended** to mix them on this model. 7. **Prefix cache**: setting `response_format` (or not setting it) **does not invalidate the prefix cache**, so you can adjust it on a per-request basis without hurting cache hit rates. ## Common errors ### `invalid_request_error` When the schema format itself is invalid (for example, `json_schema.schema` is not an object), the API returns `400` with error type `invalid_request_error`: ```json theme={null} { "error": { "message": "Invalid request: the `response_format.json_schema.schema` field in the request (expected type dict[string,interface]) is illegal...", "type": "invalid_request_error" } } ``` Please check that the schema is a valid JSON Schema object. ### Output truncated (`finish_reason="length"`) The model reached the `max_tokens` limit before outputting the complete JSON. We recommend: * Increasing `max_tokens` (e.g., 4096 or higher) * Simplifying the nesting depth of the schema * Shortening the input text length ### Field type mismatch / Markdown code block output On older models such as `kimi-k2.6`, the following may occur: * The returned `content` contains a Markdown code block (e.g., `json ... `), causing `json.loads` to fail. * Complex schemas such as `oneOf` / `$ref` are not strictly followed. Recommendations: * Use `kimi-k2.7-code` for Structured Output calls. * If you must use `kimi-k2.6`, strip Markdown markers in the business layer first, then validate the parsed result against the schema fields. # Troubleshooting Source: https://platform.kimi.ai/docs/guide/troubleshooting Troubleshoot common Kimi API issues involving accounts, billing, authentication, model parameters, output length, rate limits, and connectivity. No. The Kimi API Open Platform, Kimi Code, and Kimi Membership are independent products. Their billing models, balances/benefits, and API keys are not interchangeable: * The Kimi API Open Platform is pay-as-you-go with no subscription plan. To call the Kimi API, create an API key in the Open Platform console and use the endpoint for your region — see [Products and Plans](/docs/guide/product-plans). * Kimi Code is a separate coding product. Its API keys are not interchangeable with Open Platform keys. See the [Kimi Code documentation](https://www.kimi.com/code/docs/en/) for setup. * Kimi Membership (subscription) benefits do not convert into Open Platform balance, and Open Platform top-up balance cannot be used to purchase Kimi Membership or Kimi Code plans. Using a key from another product against the Open Platform endpoint returns 401 or 404 errors. See the checklist in "Why do I get 401, 404, or permission denied?" below. 429 is not a single cause. Check the `error.type` in the response first: * `engine_overloaded_error`: the service node is under high load (for example, peak-hour capacity pressure). Wait as indicated by `Retry-After`, reduce concurrency, and retry with exponential backoff. This error is caused by server-side capacity — topping up or upgrading your tier does not resolve it; * `rate_limit_reached_error`: an organization-level concurrency, RPM, TPM, or TPD limit was reached. Reduce request frequency, or see [Top-up and Rate Limits](/docs/pricing/limits) to upgrade your tier; * `exceeded_current_quota_error`: insufficient balance, an overdue account, or an expired voucher. Check `available_balance` with the [balance API](/docs/api/balance) and top up. Note that clients such as the OpenAI SDK retry automatically by default, so one operation may be amplified into multiple requests that consume rate-limit quota. Check the actual request count and client logs when troubleshooting. Requests interrupted by a 429 error are not charged. Check in the following order: 1. Whether the key comes from the product you are calling: Open Platform API keys and Kimi Code keys are not interchangeable; 2. Whether the key's region matches the endpoint: accounts, balances, and keys on platform.kimi.ai are isolated from other regional Kimi platforms; 3. Whether the account has available balance, and whether your voucher covers the target model; 4. Call `GET /v1/models` with the same key to confirm the target model is in the returned list; 5. Whether the model name matches your integration path: direct API calls and Codex use `kimi-k3`, while Claude Code uses the compatible alias `kimi-k3[1m]`. Follow the corresponding integration tutorial; 6. Clear stale environment variables, proxies, and old configurations in local routing tools such as CC Switch, and confirm which key and endpoint actually take effect. See also: [Error Codes](/docs/api/errors), [Use Kimi in Claude Code](/docs/guide/claude-code-kimi), [Use Kimi in Codex](/docs/guide/codex-kimi). Split the chain into two layers: the Kimi API and the third-party tool. 1. Call the Kimi API directly with the same key, endpoint, and model (see the cURL example in the [Kimi K3 quickstart](/docs/guide/kimi-k3-quickstart#basic-call)); 2. If the direct call fails, resolve balance, authentication, model permission, or request parameter issues first; 3. If the direct call succeeds but the third-party tool still fails, check the tool's logs — focus on protocol conversion, streaming responses, timeout settings, and automatic retries; 4. When using [Claude Code](/docs/guide/claude-code-kimi), [Codex](/docs/guide/codex-kimi), [OpenCode](/docs/guide/open-code), and similar tools, follow the corresponding tutorial — model names and configuration may differ; 5. Keep the client version, time of occurrence, `request_id`, actual request endpoint, and redacted logs for further troubleshooting. Note that third-party tools such as CC Switch and Trae are not maintained by the Kimi Open Platform. If the direct API call works but the tool fails, contact the tool's support channel as well. No result on the client does not mean the API request failed. When a coding tool stops waiting too early, a proxy disconnects, or a local timeout occurs, the client may stop displaying while the server-side request still completes and produces an actual call record. Check in order: 1. The HTTP status code and `request_id` of the request; 2. The `usage` field in the API response; 3. Whether the client retried automatically, spawned sub-agents, or looped tool calls; 4. The usage dashboard and billing details in the console; 5. Timeout and connection errors in the client logs. If the platform records still clearly disagree with the client records, email [api-service@moonshot.ai](mailto:api-service@moonshot.ai) with your organization ID, project, time of occurrence, `request_id`, model, client version, redacted logs, and billing details. First open the [Open Platform console](https://platform.kimi.ai/console) to review the usage dashboard and billing details. Reconcile them against client logs and the `usage` returned by the API by time, project, model, and `request_id`. If you need help from the platform team, prepare: organization ID, project name, time of occurrence (with time zone), `request_id`, model name, client and version, redacted logs, and relevant billing or export records, then email [api-service@moonshot.ai](mailto:api-service@moonshot.ai). No. The Kimi API automatically attempts to cache repeated initial context — no cache ID, TTL, or extra request parameters are needed. Keeping the initial prefix stable (system prompt, tool definitions, long documents) helps subsequent requests hit the cache. Modifying the prefix may reduce the hit rate. See [Context Caching](/docs/guide/use-context-caching-feature-of-kimi-api). Kimi K3 is unlocked after a successful top-up (minimum \$1). Your cumulative top-up amount also determines your account tier and rate limits — see [Recharge and Rate Limits](/docs/pricing/limits) and the [Kimi K3 quickstart](/docs/guide/kimi-k3-quickstart). K3 always thinks. Set the reasoning effort with the top-level `reasoning_effort` field: `low`, `high`, or `max` (default `max`). Use higher levels for more complex tasks; lower levels reduce latency and token consumption for simple tasks. See [Reasoning Effort](/docs/guide/use-reasoning-effort) for details and examples. You can't — K3 always thinks. If the reasoning takes too long, set `reasoning_effort` to `low` to reduce the reasoning effort. See [Reasoning Effort](/docs/guide/use-reasoning-effort). You can run a minimal test in the [Playground](https://platform.kimi.ai/playground) to confirm whether a model and prompt fit your scenario, and use the [MoonPalace debugging tool](/docs/guide/use-moonpalace) to capture complete requests while coding. Available models and voucher coverage are subject to what the Playground and account pages actually show. When using tool calls (`tool_calls`), the model may issue multiple consecutive tool calls based on the context. If you find that the model repeatedly calls **the same tool**, each call uses exactly the same `function.name` and `function.arguments`, and the tool result does not provide new useful information, you can treat it as a repeated tool call. When handling this issue, we recommend checking the message layout first: 1. When the Kimi API returns `finish_reason=tool_calls`, make sure the returned `choice.message` has been added to the `messages` list as is. 2. Make sure each `tool_call` has a corresponding message with `role=tool`. 3. Make sure the `tool_call_id` in the `role=tool` message exactly matches the corresponding `tool_call.id`. 4. If you use streaming output with `stream=True`, make sure the streamed `tool_calls` chunks have been assembled correctly, especially the `function.arguments` field. If the message layout is correct but the model still repeatedly calls the same tool with the same arguments, you can add repeated-call detection on the client side and append a reminder to the system prompt in the next request. When the same tool and the same arguments are repeated 3 consecutive times, you can append: ```text theme={null} You are repeating the exact same tool call with identical parameters. Please carefully analyze the previous result. If the task is not yet complete, try a different method or parameters instead of repeating the same call. ``` When the repeated call reaches 5 consecutive times, you can append a stronger reminder that includes the tool name, repeat count, and arguments: ```text theme={null} You have repeatedly called the same tool with identical parameters many times. Repeated tool call detected: - tool: {tool_name} - repeated_times: {repeat_count} - arguments: {tool_arguments} The previous repeated calls did not make progress. Do not call this exact same tool with the exact same arguments again. Carefully inspect the latest tool result and choose a different next action, different parameters, or finish the task if enough evidence has been gathered. ``` If the same tool and the same arguments are repeated 8 consecutive times, we recommend appending the stronger reminder again. Note: `` is only an example prompt, not a special field of the Kimi API. You can merge it into the next `role=system` message, or write it into the system prompt based on your own message management logic. To avoid false positives, we recommend triggering this reminder only when the same tool, the same arguments, consecutive repetition, and no new progress from the tool result are all true. The Kimi API and Kimi Assistant are different product experiences. The model version, system prompt, context management, tool configuration, and product policies may differ, so identical input does not necessarily produce identical output. When using the API, choose a model, set the system prompt, manage conversation context, and declare the required tools for your application. See the [Model List](/docs/models) and [Model Parameter Reference](/docs/api/models-overview) for available models and parameter differences. The web search (`web_search`) is currently being updated. We do not recommend using this functionality in the near term. This documentation is outdated; please follow subsequent content updates. Yes. The Kimi API provides the built-in `$web_search` tool. Declare it as a `builtin_function` in the request's `tools` field and handle its results through the standard `tool_calls` flow. Web search is not enabled automatically for every API request. See [Use Web Search with the Kimi API](/docs/guide/use-web-search) for the declaration and complete example. To connect your own or a third-party search service, see [Use the Kimi API for Tool Calls](/docs/guide/use-kimi-api-to-complete-tool-calls). If the content returned by the Kimi API is incomplete, truncated, or shorter than expected, first check the value of `choice.finish_reason` in the response body. If the value is `length`, it means the number of tokens generated by the current model exceeded the `max_completion_tokens` parameter in the request. In this case, the Kimi API only returns up to `max_completion_tokens` tokens, and any extra content is discarded. When you encounter `finish_reason=length`, and you want the Kimi model to continue from the previous response, you can use Partial Mode. For details, see: [Use Partial Mode with the Kimi API](/docs/guide/use-partial-mode-feature-of-kimi-api) To avoid `finish_reason=length`, we recommend increasing `max_completion_tokens` appropriately. A best practice is to use the [estimate-token-count](/docs/api/estimate) API to calculate the number of tokens in the input, then subtract that number from the maximum context window supported by the selected model. For example, `kimi-k3` supports up to 1M tokens, `moonshot-v1-32k` supports up to 32k tokens, while `kimi-k2.6`, `kimi-k2.5`, `kimi-k2-0905-preview`, and `kimi-k2-turbo-preview` support up to 256k tokens. The remaining value can be used as the upper bound for `max_completion_tokens` in the current request. * For `kimi-k3`, `max_completion_tokens` defaults to 131072, and the maximum output length is `1024*1024 - prompt_tokens`. * For `moonshot-v1-8k`, the maximum output length is `8*1024 - prompt_tokens`. * For `moonshot-v1-32k`, the maximum output length is `32*1024 - prompt_tokens`. * For `moonshot-v1-128k`, the maximum output length is `128*1024 - prompt_tokens`. * For `kimi-k2.6`, `kimi-k2.5`, `kimi-k2-0905-preview`, and `kimi-k2-turbo-preview`, the maximum output length is `256*1024 - prompt_tokens`. * `kimi-k3` supports approximately 1,500,000 Chinese characters. * `moonshot-v1-8k` supports approximately 15,000 Chinese characters. * `moonshot-v1-32k` supports approximately 60,000 Chinese characters. * `moonshot-v1-128k` supports approximately 200,000 Chinese characters. * `kimi-k2.6`, `kimi-k2.5`, `kimi-k2-0905-preview`, and `kimi-k2-turbo-preview` support approximately 400,000 Chinese characters. *Note: These are estimates. Actual results may vary.* We provide file upload and parsing services for various file formats. **For text files, we extract the text content. For image files, we use OCR to recognize text in the image. For PDF documents, if the PDF only contains images, we use OCR to extract text from those images; otherwise, we only extract the text content.** *Note: For images, we only use OCR to extract text. If your image does not contain any text, parsing will fail.* For the complete list of supported file formats, see: [Files Upload API](/docs/api/files-upload) We currently do not support referencing file content as context through a file `file_id`. The input to the Kimi API, or the output from the Kimi model, contains unsafe or sensitive content. **Note: Content generated by the Kimi model may also contain unsafe or sensitive content, which can trigger a `content_filter` error.** If you call the API through a third-party platform or tool, first confirm that the error was actually returned by the Kimi API: third-party platforms may apply their own content safety policies and error wording, so their messages do not necessarily come from the Kimi API. Check the third-party platform's logs as well. The platform cannot disclose the exact safety rules that were triggered. You can try narrowing the request scope, removing content that may cause false positives, and retry. If you frequently encounter errors such as `Connection Error` or `Connection Time Out` while using the Kimi API, check the following in order: 1. Whether your program code or SDK has a default timeout setting. 2. Whether you are using any type of proxy server, and whether the proxy server network and timeout settings are correct. Another possible cause is generating too many tokens without enabling streaming output with `stream=True`. In this case, the request may wait too long for generation to finish and trigger a timeout in an intermediate gateway. Some gateway applications determine whether a request is valid by checking whether the server has returned a `status_code` and `header`. When `stream=True` is not used, the Kimi server waits until generation is complete before sending the `header`. While waiting for that `header`, some gateway applications may close long-running connections, causing connection-related errors. **We recommend enabling streaming output with `stream=True` to reduce connection-related errors as much as possible.** If you encounter a `rate_limit_reached_error` while using the Kimi API, for example: ```text theme={null} rate_limit_reached_error: Your account {uid}<{ak-id}> request reached TPM rate limit, current:{current_tpm}, limit:{max_tpm} ``` but the TPM or RPM limit in the error message does not match the TPM or RPM shown in the dashboard, first check whether you are using the correct `api_key` for the current account. In most cases, this mismatch is caused by using the wrong `api_key`, such as accidentally using a key provided by another user or mixing up keys across multiple accounts. Make sure your SDK is configured with `base_url=https://api.moonshot.ai/v1`. The `model_not_found` error is usually caused by using the OpenAI SDK without setting `base_url`, which sends the request to OpenAI servers instead. OpenAI then returns the `model_not_found` error. Due to the uncertainty of model generation, the Kimi model may make calculation errors of varying severity when performing numerical computations. We recommend using tool calls (`tool_calls`) to provide calculator functionality to the Kimi model. For more information, see our tool calling guide: [Use the Kimi API for Tool Calls (`tool_calls`)](/docs/guide/use-kimi-api-to-complete-tool-calls) The Kimi model cannot access highly time-sensitive information such as the current date. However, you can provide this information in the system prompt. For example: The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ```python theme={null} import os from datetime import datetime from openai import OpenAI client = OpenAI( api_key=os.environ['MOONSHOT_API_KEY'], base_url="https://api.moonshot.ai/v1", ) # Generate the current date with datetime and add it to the system prompt. system_prompt = f""" You are Kimi, and today's date is {datetime.now().strftime('%d.%m.%Y %H:%M:%S')} """ completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "system", "content": system_prompt}, {"role": "user", "content": "What is today's date?"}, ], ) print(completion.choices[0].message.content) # Output: Today's date is July 31, 2024. ``` ```js theme={null} const OpenAI = require('openai') client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }) // Generate the current date with Date and add it to the system prompt. system_prompt = `You are Kimi, and today's date is ${new Date().toString()}` async function main() { completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ {role: "system", content: system_prompt}, {role: "user", content: "What is today's date?"}, ], }) console.log(completion.choices[0].message.content) // Output: Today's date is July 31, 2024. } main() ``` In some cases, you may need to integrate with the Kimi API directly instead of using the OpenAI SDK. When integrating directly, decide the next processing step based on the status returned by the API. In general, HTTP status code 200 indicates success, while 4xx and 5xx status codes indicate failure. We return error information in JSON format. See the following code snippets for example handling logic: ```python theme={null} import os import httpx header = { "Authorization": f"Bearer {os.environ['MOONSHOT_API_KEY']}", } messages = [ {"role": "system", "content": "You are Kimi"}, {"role": "user", "content": "Hello."}, ] r = httpx.post("https://api.moonshot.ai/v1/chat/completions", headers=header, json={ "model": "kimi-k3", # A correct model enters the status_code == 200 branch below. # "model": "moonshot-v1-129k", # An incorrect model name enters the else branch below. "messages": messages, }) if r.status_code == 200: # Normal handling for a request with a correct model. completion = r.json() print(completion["choices"][0]["message"]["content"]) else: # Error handling for a request with an incorrect model name. # For demonstration, we only print the error here. # In production logic, you may want to log, interrupt the request, retry, and so on. error = r.json() print(f"error: status={r.status_code}, type='{error['error']['type']}', message='{error['error']['message']}'") ``` ```js theme={null} const axios = require('axios'); header = { "Authorization": `Bearer ${process.env.MOONSHOT_API_KEY}`, } messages = [ {"role": "system", "content": "You are Kimi"}, {"role": "user", "content": "Hello."}, ] async function main() { r = await axios.post("https://api.moonshot.ai/v1/chat/completions", { "model": "kimi-k3", // A correct model enters the success branch below. //"model": "moonshot-v1-129k", // An incorrect model name enters the error branch. "messages": messages, }, { headers: header, validateStatus: function (status) { return status == 200; // Resolve only when the status code is 200. } }, ).catch(function (error) { console.log(`error: ${error.message}`) }) if (r) { // Normal handling for a request with a correct model. console.log(r.data.choices[0].message.content) } } main() ``` Our error messages follow this format: ```json theme={null} { "error": { "type": "error_type", "message": "error_message" } } ``` For the complete error reference, see: [Error Reference](/docs/api/errors) If some requests with similar prompts respond quickly, for example in 3 seconds, while others respond slowly, for example in 20 seconds, this is usually because the Kimi model generated different numbers of tokens. In general, the number of generated tokens is proportional to the total response time of the Kimi API. The more tokens generated, the longer the full response takes. Note that the number of generated tokens only affects the response time for the complete request, meaning the time until the final token is generated. You can set `stream=True` and observe the time to first token, abbreviated as TTFT. In normal cases, when prompt lengths are similar, TTFT should not vary significantly. > Note: `max_tokens` is deprecated. Please use `max_completion_tokens` instead. The two fields have the same meaning. The `max_completion_tokens` parameter means: **when calling `/v1/chat/completions`, it specifies the maximum number of tokens the model is allowed to generate. Once the number of generated tokens exceeds `max_completion_tokens`, the model stops generating the next token**. `max_completion_tokens` is used to: 1. Help the caller determine which model to use. For example, when `prompt_tokens + max_completion_tokens <= 8 * 1024`, you can choose `moonshot-v1-8k`. 2. Prevent the Kimi model from generating too much unexpected content in unusual cases, which could cause extra cost. For example, the model might repeatedly output whitespace. `max_completion_tokens` does not tell the Kimi model how many tokens to output. In other words, **`max_completion_tokens` is not used as part of the prompt input to the Kimi model**. If you want the model to output a specific number of characters, use these general approaches: * For outputs under 1,000 characters: 1. Clearly specify the desired character count in the prompt. 2. Manually or programmatically check whether the output length meets expectations. If not, tell the model in a second round that the output is too long or too short, then ask it to generate a new version. * For outputs over 1,000 characters or much longer: 1. Split the expected output by structure or chapter, create a template, and use placeholders to mark where the model should fill in content. 2. Ask the Kimi model to fill each placeholder one by one, then assemble the complete long-form text. The OpenAI SDK usually includes a retry mechanism: > Certain errors are automatically retried 2 times by default, with a short exponential backoff. Connection errors (for example, due to a network connectivity problem), 408 Request Timeout, 409 Conflict, 429 Rate Limit, and >=500 Internal errors are all retried by default. This retry mechanism retries 2 times by default when an error occurs, for a total of 3 requests. In unstable network conditions or other cases that may cause request errors, using the OpenAI SDK can expand one request into 2 or 3 requests. All of these requests count toward your RPM quota. *Note: For users on a `tier0` account using the OpenAI SDK, one failed request may consume the entire RPM quota because of the default retry mechanism.* Please do not do this. Encoding files with `base64` can cause massive token consumption. If your file type is supported by our `/v1/files` API, upload the file through the files API and extract its content there. For binary files or files in other encoded formats, the Kimi model currently cannot parse the content. Do not add this content to the context. Kimi Open Platform provides separate regional platforms. For users outside mainland China, use platform.kimi.ai. Accounts and keys from different regional platforms are independent and cannot be mixed. If you use the wrong key for a platform, you will receive a `401 invalid_authentication_error`. If you receive a 401 error, first check whether you are using a key from the wrong platform. * International open platform `base_url`: [https://api.moonshot.ai/v1](https://api.moonshot.ai/v1) # Using Batch API for Bulk Processing Source: https://platform.kimi.ai/docs/guide/use-batch-api Upload JSONL requests, create and monitor Batch API jobs, and download results for large-scale asynchronous inference. When you need to process large-scale tasks with low real-time requirements, the Batch API is the ideal choice. It supports submitting tasks in bulk via files, saving 40% on inference costs compared to real-time API calls. Batch API supports both the `kimi-k2.6` and `kimi-k2.5` models; `kimi-k3` is not supported. The `temperature`, `top_p`, and other parameters cannot be modified for these models — do not include them in the request body. Upload a JSONL file and create a batch task List batch tasks for your organization Get status and details for a specific batch task Cancel an in-progress batch task ## Workflow This guide walks through a complete text classification example using the Batch API: ### 1. Build the Input File Each line in the JSONL file is an independent JSON object representing a single inference request: ```json theme={null} {"custom_id": "request-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "kimi-k2.6", "messages": [{"role": "system", "content": "You are a text classification assistant."}, {"role": "user", "content": "Classify this text: AI is transforming the world"}]}} ``` | Field | Required | Description | | ----------- | -------- | ------------------------------------------------------------------------------ | | `custom_id` | Yes | Custom request identifier for tracking results, must be unique within the file | | `method` | Yes | Request method, must be `POST` | | `url` | Yes | Request endpoint, must be `/v1/chat/completions` | | `body` | Yes | Request body, same parameters as the [Chat Completions API](/docs/api/chat) | The `model` in `body` must be either `kimi-k2.6` or `kimi-k2.5`. The `temperature`, `top_p`, `n`, `presence_penalty`, and `frequency_penalty` parameters cannot be modified for these models. Do not include these parameters in the `body`. **Input file requirements:** * File must be in `.jsonl` format, non-empty, and no larger than 100MB * Each line must be a valid JSON object containing `custom_id`, `method`, `url`, and `body` fields * `custom_id` must be unique within the file * All lines must use the same `model` — only one model per batch is allowed * `method` must be `POST`, `url` must be `/v1/chat/completions` * The specified model must exist and the user must have access to it ### 2. Upload the File Upload the JSONL file via the [Upload File](/docs/api/files-upload) endpoint with `purpose` set to `"batch"`. ```python Python theme={null} import os from openai import OpenAI from openai.types import FileObject client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) file_object: FileObject = client.files.create( file=open("batch_requests.jsonl", "rb"), purpose="batch", ) print(file_object.id) # Save file_id for the next step ``` ```bash cURL theme={null} curl ${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/files \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -F purpose="batch" \ -F file="@batch_requests.jsonl" ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const fs = require("fs"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { const fileObject = await client.files.create({ file: fs.createReadStream("batch_requests.jsonl"), purpose: "batch" }); console.log(fileObject.id); // Save file_id for the next step } main(); ``` ### 3. Create the Task Call the [Create Batch](/docs/api/batch-create) endpoint with `input_file_id` and `completion_window`. We recommend setting a generous time window for larger datasets. ```python Python theme={null} import os from openai import OpenAI from openai.types import Batch client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) batch: Batch = client.batches.create( input_file_id="your_file_id", endpoint="/v1/chat/completions", completion_window="24h", ) print(batch.id) # Save batch_id for polling ``` ```bash cURL theme={null} curl ${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/batches \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "input_file_id": "your_file_id", "endpoint": "/v1/chat/completions", "completion_window": "24h" }' ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { const batch = await client.batches.create({ input_file_id: "your_file_id", endpoint: "/v1/chat/completions", completion_window: "24h" }); console.log(batch.id); // Save batch_id for polling } main(); ``` ### 4. Wait for Completion After creation, the task enters `validating` status for input validation. Once validated, it moves to `in_progress`. Use the [Retrieve Batch](/docs/api/batch-retrieve) endpoint to poll for status updates. ```python Python theme={null} import os import time from openai import OpenAI from openai.types import Batch client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) while True: batch: Batch = client.batches.retrieve("your_batch_id") completed: int = batch.request_counts.completed if batch.request_counts else 0 total: int = batch.request_counts.total if batch.request_counts else 0 print(f"Status: {batch.status} ({completed}/{total})") if batch.status == "completed": break elif batch.status in ("failed", "expired", "cancelled"): print(f"Task terminated: {batch.status}") break time.sleep(10) ``` ```bash cURL theme={null} curl ${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/batches/your_batch_id \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { let batch = await client.batches.retrieve("your_batch_id"); while (!["completed", "failed", "expired", "cancelled"].includes(batch.status)) { await new Promise(r => setTimeout(r, 10000)); batch = await client.batches.retrieve("your_batch_id"); console.log(`Status: ${batch.status} (${batch.request_counts.completed}/${batch.request_counts.total})`); } } main(); ``` ### 5. Process Results When complete, `output_file_id` contains the results file ID. Download it via the [Get File Content](/docs/api/files-content) endpoint. If any requests failed, `error_file_id` contains the error details. ```python Python theme={null} import json import os from openai import OpenAI client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) output = client.files.content("your_output_file_id") for line in output.text.strip().split("\n"): result: dict = json.loads(line) custom_id: str = result["custom_id"] content: str = result["response"]["body"]["choices"][0]["message"]["content"] print(f"{custom_id}: {content}") ``` ```bash cURL theme={null} curl ${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/files/your_output_file_id/content \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -o results.jsonl ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { const output = await client.files.content("your_output_file_id"); const text = await output.text(); for (const line of text.trim().split("\n")) { const data = JSON.parse(line); console.log(`${data.custom_id}: ${data.response.body.choices[0].message.content}`); } } main(); ``` Each line in the output file corresponds to a processed request: ```json theme={null} { "id": "request-1", "custom_id": "request-1", "response": { "status_code": 200, "request_id": "", "body": { "id": "chatcmpl-xxx", "object": "chat.completion", "created": 1711475054, "model": "kimi-k2.6", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "This text belongs to the Technology category." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 30, "completion_tokens": 10, "total_tokens": 40 } } }, "error": null } ``` ## Complete Code Examples Complete scripts combining all steps above — copy and run directly: ```python Python expandable theme={null} import json import os import time from pathlib import Path from openai import OpenAI MODEL = "kimi-k2.6" client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) def create_input_jsonl() -> Path: """Build a JSONL input file with classification requests.""" texts: list[str] = [ "Hamlet is one of Shakespeare's most famous tragedies", "Scientists discover new potentially habitable planet", "2024 Artificial Intelligence Development Report", "How to make a delicious braised pork dish", "Latest iPhone launch event details", ] requests: list[dict] = [] for i, text in enumerate(texts): requests.append({ "custom_id": f"text_{i}", "method": "POST", "url": "/v1/chat/completions", "body": { "model": MODEL, "messages": [ {"role": "system", "content": "You are a text classification expert. Classify texts into: Literature/News/Academic/Technology/Lifestyle"}, {"role": "user", "content": f"Please classify the following text: {text}"}, ], }, }) output_path = Path("classification_requests.jsonl") with output_path.open("w", encoding="utf-8") as f: for req in requests: f.write(json.dumps(req, ensure_ascii=False) + "\n") return output_path # 1. Build input file input_file: Path = create_input_jsonl() # 2. Upload file file_object = client.files.create(file=input_file, purpose="batch") print(f"File uploaded: {file_object.id}") # 3. Create batch task batch = client.batches.create( input_file_id=file_object.id, endpoint="/v1/chat/completions", completion_window="24h", ) print(f"Batch created: {batch.id}") # 4. Poll for completion while True: batch = client.batches.retrieve(batch.id) print(f"Status: {batch.status} ({batch.request_counts.completed}/{batch.request_counts.total})") if batch.status == "completed": break elif batch.status in ("failed", "expired", "cancelled"): print(f"Task terminated: {batch.status}") exit(1) time.sleep(10) # 5. Process results output = client.files.content(batch.output_file_id) for line in output.text.strip().split("\n"): data: dict = json.loads(line) print(f"{data['custom_id']}: {data['response']['body']['choices'][0]['message']['content']}") ``` ```javascript Node.js expandable theme={null} const OpenAI = require("openai"); const fs = require("fs"); const MODEL = "kimi-k2.6"; const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); // 1. Build input file const texts = [ "Hamlet is one of Shakespeare's most famous tragedies", "Scientists discover new potentially habitable planet", "2024 Artificial Intelligence Development Report", "How to make a delicious braised pork dish", "Latest iPhone launch event details" ]; const lines = texts.map((text, i) => JSON.stringify({ custom_id: `text_${i}`, method: "POST", url: "/v1/chat/completions", body: { model: MODEL, messages: [ { role: "system", content: "You are a text classification expert. Classify texts into: Literature/News/Academic/Technology/Lifestyle" }, { role: "user", content: `Please classify the following text: ${text}` } ] } })); fs.writeFileSync("classification_requests.jsonl", lines.join("\n") + "\n"); async function main() { // 2. Upload file const fileObject = await client.files.create({ file: fs.createReadStream("classification_requests.jsonl"), purpose: "batch" }); console.log(`File uploaded: ${fileObject.id}`); // 3. Create batch task const batch = await client.batches.create({ input_file_id: fileObject.id, endpoint: "/v1/chat/completions", completion_window: "24h" }); console.log(`Batch created: ${batch.id}`); // 4. Poll for completion let current = batch; while (!["completed", "failed", "expired", "cancelled"].includes(current.status)) { await new Promise(r => setTimeout(r, 10000)); current = await client.batches.retrieve(batch.id); console.log(`Status: ${current.status} (${current.request_counts.completed}/${current.request_counts.total})`); } if (current.status !== "completed") { console.error(`Task terminated: ${current.status}`); return; } // 5. Download and process results const output = await client.files.content(current.output_file_id); const text = await output.text(); for (const line of text.trim().split("\n")) { const data = JSON.parse(line); console.log(`${data.custom_id}: ${data.response.body.choices[0].message.content}`); } } main(); ``` ## Batch Status Reference | Status | Description | | ------------- | ------------------------------------------- | | `validating` | Created, input data validation in progress | | `failed` | Data validation failed, batch terminated | | `in_progress` | Validation passed, execution in progress | | `finalizing` | Execution complete, preparing results | | `completed` | Results ready, batch complete | | `expired` | Did not complete within `completion_window` | | `cancelling` | Cancellation requested, pending | | `cancelled` | Cancellation complete, batch terminated | ## Task Management ### List Batches Use the [List Batches](/docs/api/batch-list) endpoint to view all batch tasks in your organization. ```python Python theme={null} import os from openai import OpenAI from openai.pagination import SyncCursorPage from openai.types import Batch client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) batches: SyncCursorPage[Batch] = client.batches.list(limit=10) for batch in batches.data: print(f"{batch.id} - {batch.status} ({batch.request_counts.completed}/{batch.request_counts.total})") ``` ```bash cURL theme={null} curl "${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/batches?limit=10" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { const batches = await client.batches.list({ limit: 10 }); for (const batch of batches.data) { console.log(`${batch.id} - ${batch.status} (${batch.request_counts.completed}/${batch.request_counts.total})`); } } main(); ``` ### Cancel a Batch Use the [Cancel Batch](/docs/api/batch-cancel) endpoint to cancel an in-progress task. Only tasks in `validating`, `in_progress`, or `finalizing` status can be cancelled. After cancellation, the status changes to `cancelling` and then `cancelled`. ```python Python theme={null} import os from openai import OpenAI from openai.types import Batch client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) batch: Batch = client.batches.cancel("your_batch_id") print(f"Status: {batch.status}") # cancelling ``` ```bash cURL theme={null} curl -X POST ${MOONSHOT_BASE_URL:-https://api.moonshot.ai/v1}/batches/your_batch_id/cancel \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` ```javascript Node.js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); async function main() { const batch = await client.batches.cancel("your_batch_id"); console.log(`Status: ${batch.status}`); // cancelling } main(); ``` ## Multi-modal Batch Tasks The Batch API supports image and video content in the input file. The key difference from text tasks is in **building the input file** — the rest of the workflow (upload, create task, poll, process results) is identical. There are two ways to include images: * **Base64 inline**: Encode images as base64 directly in the JSONL. Suitable for small images. Note that base64 inflates file size by \~33% — keep the 100MB file size limit in mind. * **File reference**: Upload images first via the Files API (`purpose="image"`), then reference them in the JSONL using `ms://`. Better for large images or image reuse. Both methods are provided below — switch between them as needed: ```python Python expandable theme={null} import base64 import json import os import time from pathlib import Path from openai import OpenAI from openai.types import Batch, FileObject client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) MODEL = "kimi-k2.6" PROMPT = "Classify this image: Landscape/Portrait/Food/Architecture/Other" SYSTEM = "You are an image classification assistant." def build_request_base64(custom_id: str, image_path: str) -> dict: """Method 1: Encode the image as base64 and embed it directly in the JSONL. Best for small images — no extra upload step needed.""" with open(image_path, "rb") as f: image_data: str = base64.b64encode(f.read()).decode("utf-8") return { "custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": { "model": MODEL, "messages": [ {"role": "system", "content": SYSTEM}, { "role": "user", "content": [ {"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image_data}"}}, {"type": "text", "text": PROMPT}, ], }, ], }, } def build_request_upload(custom_id: str, image_path: str) -> dict: """Method 2: Upload the image first, then reference it via ms://. Best for large images or when the same image is reused across requests.""" file_object: FileObject = client.files.create( file=open(image_path, "rb"), purpose="image", ) print(f"Image uploaded: {image_path} -> {file_object.id}") return { "custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": { "model": MODEL, "messages": [ {"role": "system", "content": SYSTEM}, { "role": "user", "content": [ {"type": "image_url", "image_url": {"url": f"ms://{file_object.id}"}}, {"type": "text", "text": PROMPT}, ], }, ], }, } # ====== Choose your build method here ====== build_request = build_request_base64 # or build_request_upload # ============================================ # 1. Build input file images: list[str] = ["image1.png", "image2.png", "image3.png"] requests: list[dict] = [build_request(f"img-{i}", path) for i, path in enumerate(images)] input_path = Path("image_batch_requests.jsonl") with input_path.open("w", encoding="utf-8") as f: for req in requests: f.write(json.dumps(req, ensure_ascii=False) + "\n") # 2. Upload JSONL and create task file_object: FileObject = client.files.create(file=input_path, purpose="batch") batch: Batch = client.batches.create( input_file_id=file_object.id, endpoint="/v1/chat/completions", completion_window="24h", ) print(f"Batch created: {batch.id}") # 3. Poll for completion while True: batch = client.batches.retrieve(batch.id) print(f"Status: {batch.status} ({batch.request_counts.completed}/{batch.request_counts.total})") if batch.status == "completed": break elif batch.status in ("failed", "expired", "cancelled"): print(f"Task terminated: {batch.status}") exit(1) time.sleep(10) # 4. Process results output = client.files.content(batch.output_file_id) for line in output.text.strip().split("\n"): data: dict = json.loads(line) print(f"{data['custom_id']}: {data['response']['body']['choices'][0]['message']['content']}") ``` ```javascript Node.js expandable theme={null} const OpenAI = require("openai"); const fs = require("fs"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); const MODEL = "kimi-k2.6"; const PROMPT = "Classify this image: Landscape/Portrait/Food/Architecture/Other"; const SYSTEM = "You are an image classification assistant."; /** Method 1: Encode the image as base64 and embed it directly in the JSONL. * Best for small images — no extra upload step needed. */ function buildRequestBase64(customId, imagePath) { const imageData = fs.readFileSync(imagePath).toString("base64"); return { custom_id: customId, method: "POST", url: "/v1/chat/completions", body: { model: MODEL, messages: [ { role: "system", content: SYSTEM }, { role: "user", content: [ { type: "image_url", image_url: { url: `data:image/png;base64,${imageData}` } }, { type: "text", text: PROMPT }, ], }, ], }, }; } /** Method 2: Upload the image first, then reference it via ms://. * Best for large images or when the same image is reused across requests. */ async function buildRequestUpload(customId, imagePath) { const fileObject = await client.files.create({ file: fs.createReadStream(imagePath), purpose: "image" }); console.log(`Image uploaded: ${imagePath} -> ${fileObject.id}`); return { custom_id: customId, method: "POST", url: "/v1/chat/completions", body: { model: MODEL, messages: [ { role: "system", content: SYSTEM }, { role: "user", content: [ { type: "image_url", image_url: { url: `ms://${fileObject.id}` } }, { type: "text", text: PROMPT }, ], }, ], }, }; } async function main() { // ====== Choose your build method here ====== const useUpload = false; // set to true for file reference method // ============================================ // 1. Build input file const images = ["image1.png", "image2.png", "image3.png"]; const requests = []; for (let i = 0; i < images.length; i++) { const req = useUpload ? await buildRequestUpload(`img-${i}`, images[i]) : buildRequestBase64(`img-${i}`, images[i]); requests.push(JSON.stringify(req)); } fs.writeFileSync("image_batch_requests.jsonl", requests.join("\n") + "\n"); // 2. Upload JSONL and create task const fileObject = await client.files.create({ file: fs.createReadStream("image_batch_requests.jsonl"), purpose: "batch" }); const batch = await client.batches.create({ input_file_id: fileObject.id, endpoint: "/v1/chat/completions", completion_window: "24h" }); console.log(`Batch created: ${batch.id}`); // 3. Poll for completion let current = batch; while (!["completed", "failed", "expired", "cancelled"].includes(current.status)) { await new Promise(r => setTimeout(r, 10000)); current = await client.batches.retrieve(batch.id); console.log(`Status: ${current.status} (${current.request_counts.completed}/${current.request_counts.total})`); } if (current.status !== "completed") { console.error(`Task terminated: ${current.status}`); return; } // 4. Process results const output = await client.files.content(current.output_file_id); const text = await output.text(); for (const line of text.trim().split("\n")) { const data = JSON.parse(line); console.log(`${data.custom_id}: ${data.response.body.choices[0].message.content}`); } } main(); ``` There are two ways to include videos: * **Base64 inline**: Encode videos as base64 directly in the JSONL. Suitable for small videos. Note that base64 inflates file size by \~33% — keep the 100MB file size limit in mind. * **File reference**: Upload videos first via the Files API (`purpose="video"`), then reference them in the JSONL using `ms://`. Better for large videos or video reuse. Both methods are provided below — switch between them as needed: ```python Python expandable theme={null} import base64 import json import os import time from pathlib import Path from openai import OpenAI from openai.types import Batch, FileObject MODEL = "kimi-k2.6" client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url=os.environ.get("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1"), ) PROMPT = "Summarize the main content of this video." SYSTEM = "You are a video content analysis assistant." def build_request_base64(custom_id: str, video_path: str) -> dict: """Method 1: Encode the video as base64 and embed it directly in the JSONL. Best for small videos — no extra upload step needed.""" with open(video_path, "rb") as f: video_data: str = base64.b64encode(f.read()).decode("utf-8") return { "custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": { "model": MODEL, "messages": [ {"role": "system", "content": SYSTEM}, { "role": "user", "content": [ {"type": "video_url", "video_url": {"url": f"data:video/mp4;base64,{video_data}"}}, {"type": "text", "text": PROMPT}, ], }, ], }, } def build_request_upload(custom_id: str, video_path: str) -> dict: """Method 2: Upload the video first, then reference it via ms://. Best for large videos or when the same video is reused across requests.""" file_object: FileObject = client.files.create( file=open(video_path, "rb"), purpose="video", ) print(f"Video uploaded: {video_path} -> {file_object.id}") return { "custom_id": custom_id, "method": "POST", "url": "/v1/chat/completions", "body": { "model": MODEL, "messages": [ {"role": "system", "content": SYSTEM}, { "role": "user", "content": [ {"type": "video_url", "video_url": {"url": f"ms://{file_object.id}"}}, {"type": "text", "text": PROMPT}, ], }, ], }, } # ====== Choose your build method here ====== build_request = build_request_base64 # or build_request_upload # ============================================ # 1. Build input file videos: list[str] = ["video1.mp4", "video2.mp4", "video3.mp4"] requests: list[dict] = [build_request(f"video-{i}", path) for i, path in enumerate(videos)] input_path = Path("video_batch_requests.jsonl") with input_path.open("w", encoding="utf-8") as f: for req in requests: f.write(json.dumps(req, ensure_ascii=False) + "\n") # 2. Upload JSONL and create task batch_file: FileObject = client.files.create(file=input_path, purpose="batch") batch: Batch = client.batches.create( input_file_id=batch_file.id, endpoint="/v1/chat/completions", completion_window="24h", ) print(f"Batch created: {batch.id}") # 3. Poll for completion while True: batch = client.batches.retrieve(batch.id) print(f"Status: {batch.status} ({batch.request_counts.completed}/{batch.request_counts.total})") if batch.status == "completed": break elif batch.status in ("failed", "expired", "cancelled"): print(f"Task terminated: {batch.status}") exit(1) time.sleep(10) # 4. Process results output = client.files.content(batch.output_file_id) for line in output.text.strip().split("\n"): data: dict = json.loads(line) print(f"{data['custom_id']}: {data['response']['body']['choices'][0]['message']['content']}") ``` ```javascript Node.js expandable theme={null} const OpenAI = require("openai"); const fs = require("fs"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: process.env.MOONSHOT_BASE_URL || "https://api.moonshot.ai/v1", }); const MODEL = "kimi-k2.6"; const PROMPT = "Summarize the main content of this video."; const SYSTEM = "You are a video content analysis assistant."; /** Method 1: Encode the video as base64 and embed it directly in the JSONL. * Best for small videos — no extra upload step needed. */ function buildRequestBase64(customId, videoPath) { const videoData = fs.readFileSync(videoPath).toString("base64"); return { custom_id: customId, method: "POST", url: "/v1/chat/completions", body: { model: MODEL, messages: [ { role: "system", content: SYSTEM }, { role: "user", content: [ { type: "video_url", video_url: { url: `data:video/mp4;base64,${videoData}` } }, { type: "text", text: PROMPT }, ], }, ], }, }; } /** Method 2: Upload the video first, then reference it via ms://. * Best for large videos or when the same video is reused across requests. */ async function buildRequestUpload(customId, videoPath) { const fileObject = await client.files.create({ file: fs.createReadStream(videoPath), purpose: "video" }); console.log(`Video uploaded: ${videoPath} -> ${fileObject.id}`); return { custom_id: customId, method: "POST", url: "/v1/chat/completions", body: { model: MODEL, messages: [ { role: "system", content: SYSTEM }, { role: "user", content: [ { type: "video_url", video_url: { url: `ms://${fileObject.id}` } }, { type: "text", text: PROMPT }, ], }, ], }, }; } async function main() { // ====== Choose your build method here ====== const useUpload = false; // set to true for file reference method // ============================================ // 1. Build input file const videos = ["video1.mp4", "video2.mp4", "video3.mp4"]; const requests = []; for (let i = 0; i < videos.length; i++) { const req = useUpload ? await buildRequestUpload(`video-${i}`, videos[i]) : buildRequestBase64(`video-${i}`, videos[i]); requests.push(JSON.stringify(req)); } fs.writeFileSync("video_batch_requests.jsonl", requests.join("\n") + "\n"); // 2. Upload JSONL and create task const batchFile = await client.files.create({ file: fs.createReadStream("video_batch_requests.jsonl"), purpose: "batch" }); const batch = await client.batches.create({ input_file_id: batchFile.id, endpoint: "/v1/chat/completions", completion_window: "24h" }); console.log(`Batch created: ${batch.id}`); // 3. Poll for completion let current = batch; while (!["completed", "failed", "expired", "cancelled"].includes(current.status)) { await new Promise(r => setTimeout(r, 10000)); current = await client.batches.retrieve(batch.id); console.log(`Status: ${current.status} (${current.request_counts.completed}/${current.request_counts.total})`); } if (current.status !== "completed") { console.error(`Task terminated: ${current.status}`); return; } // 4. Process results const output = await client.files.content(current.output_file_id); const text = await output.text(); for (const line of text.trim().split("\n")) { const data = JSON.parse(line); console.log(`${data.custom_id}: ${data.response.body.choices[0].message.content}`); } } main(); ``` ## Best Practices * Set `completion_window` based on data volume — use `3d` or `7d` for larger datasets * Poll every 10-60 seconds to avoid excessive requests * Process results into databases or reports as needed * For very large files, consider splitting into multiple batches # Using the Console for Batch Inference Source: https://platform.kimi.ai/docs/guide/use-batch-inference Create, monitor, and download batch inference results from the Kimi Open Platform console without writing code. Batch inference allows you to submit large-scale inference tasks through the Kimi Platform console without writing any code. This tutorial shows how to create, monitor, and retrieve results of batch inference tasks from the console. Batch inference is conducted for a specific project and is only available to users at Tier1 or above. If you prefer to use the API for batch processing, see the [Batch API Guide](/docs/guide/use-batch-api). ## Steps ### 1. Create a Batch Task Open the [Kimi Platform](https://platform.kimi.ai), go to **User Center** → **Projects** → **View Project** → **Batches** → **Create Batches**. Project list Batches page ### 2. Configure Task Parameters In the Create Batch Task dialog, set the following: * **Batch Name**: Set a name for the task * **Completion Window**: Select the maximum wait time for the task * **Data File**: Upload a new file or select an existing file Click **OK** to submit the task. Create Batch Task ### 3. Wait for Inference to Complete After submission, the task starts running. You can monitor the task status in the Batches list. Task validating ### 4. Download Output Results After execution is completed, click **Details** to view the task details and download the output file. Task completed Batch details ### 5. View File History You may also find your previously uploaded files and output results in the project's **Files** page. Files list # Use the Context Caching Feature of Kimi API Source: https://platform.kimi.ai/docs/guide/use-context-caching-feature-of-kimi-api Understand Kimi API automatic context caching, cache-hit conditions, billing, usage fields, and use cases for reducing cost and latency. Context Caching pre-stores large amounts of data that may be requested frequently; when the same information is requested again, the system serves it directly from the cache instead of recomputing or retrieving it from the original source, saving time and resources. In the Kimi API, Context Caching is automatically enabled for all model requests: when the system detects repeated initial contexts (such as system prompts, knowledge documents, or tool definitions), it automatically reuses the cached content for cost optimization and faster responses, with no manual cache creation or management required. ## Use Context Caching for frequent requests over a fixed context Context Caching is especially suitable for scenarios with frequent requests that repeatedly reference a large initial context, such as: * QA bots that provide extensive preset content, such as product documentation assistants. * Frequent queries against a fixed document collection, such as public disclosure Q\&A tools for listed companies. * Periodic analysis of static codebases or knowledge bases, such as various Copilot Agents. * Viral AI applications with sudden traffic spikes. * Agent applications with complex interaction rules. ## Context Caching vs. RAG: how to choose RAG (Retrieval-Augmented Generation) is widely used in the industry for cost reduction in long-text scenarios. Context Caching's cost reduction is highly dependent on business characteristics, while RAG's is not. The main differences are: | Dimension | Context Caching | RAG | | ---------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | | Business cost | Extremely high cost compression in specific scenarios, up to 90% | Any business can reduce costs, but recall accuracy issues may degrade answer quality | | Development cost | Relatively low; the system handles caching automatically, with no additional integration or tuning needed | Relatively high; requires combining RAG with Embedding and continuous business-specific tuning | | Extra benefit | Average first-token latency can drop to within 5s in long-text scenarios | Original text length can be extended to very long, friendly to scenarios requiring millions of words of context in one go | > **Recommendation**: For frequent queries against fixed content (e.g., FAQs, document Q\&A), prioritize Context Caching; if the content is extremely long and query directions are unfixed, consider a RAG solution. ## No configuration needed: caching is automatic Context Caching uses a fully automatic caching mechanism — just call the API as usual: * **No manual creation**: The system automatically identifies and caches frequently used initial contexts. * **No cache ID references**: When calling `/v1/chat/completions`, simply pass messages in the normal way, and the system will automatically match caches in the background. * **No TTL management**: Cache lifecycle is managed automatically by the system, with no manual intervention required. The system will automatically trigger cache optimization at the appropriate times. A new request can hit the prefix cache only when the previous request's prompt tokens exceed 256. If the previous request's prompt tokens are below 256, the request is not cached and is discarded. ## Billing For Context Caching billing methods and pricing details, see the [billing information on the Product Pricing page](/docs/pricing/chat#billing-logic). ## Notes * **Cache hit conditions**: The system automatically optimizes caching for frequently repeated initial contexts. Make sure your knowledge content, system prompts, and tool definitions are relatively stable for better cache hit rates. * **Multi-turn conversations**: Place fixed large contexts (such as knowledge documents) at the beginning of the `messages` array (before the system message), then append user questions and model replies; the system will automatically identify and cache this fixed content. * **No extra configuration**: Context Caching is automatically effective for all requests — you do not need to modify your API calling method or add extra parameters; just focus on your prompt design and business logic. # Dynamically Loaded Tools Source: https://platform.kimi.ai/docs/guide/use-dynamic-tool-loading Append tool definitions to Kimi conversations on demand to reduce token usage, improve tool selection, and preserve prefix caching. When your application needs a large number of tools, declaring all of them up front in the top-level `tools` field of every request leads to **Tool Definition Bloat**: every request carries the descriptions and parameter schemas of all tools, driving up token usage, and the more candidate tools there are, the more likely the model picks the wrong tool or constructs invalid call arguments. Dynamically Loaded Tools let you **inject tools on demand** during a conversation: start with only a few core tools, and insert additional tools into `messages` when the conversation actually needs them, which reduces token usage and improves tool-selection accuracy at the same time. Because tool declarations are only ever **appended** to the end of `messages`, the existing conversation prefix stays unchanged, so dynamic loading does not break the prefix cache you have already built up and can be combined with [Context Caching](/docs/guide/use-context-caching-feature-of-kimi-api) to further reduce cost and latency. For the reasoning behind this design (lazy loading, tool registry) and combined practices, see [Kimi K3 API Tool Calling Best Practices](/docs/guide/kimi-k3-tool-calling-best-practice). Dynamically loaded tools at a glance: fetch only the tools you need, when you need them ## Inject tool declarations into messages Insert a message with `role` set to `system` into `messages`, and declare the tools to load through that message's `tools` field. The declaration format is identical to the top-level `tools` field of the request, and must contain the **complete** tool definition (`name`, `description`, `parameters`): ```json theme={null} { "messages": [ { "role": "system", "content": "You are Kimi, an AI assistant developed by Moonshot AI.\nYou are capable of a wide range of tasks, including:\n📝 Q&A & Explanation – Answer all kinds of questions, ranging from scientific knowledge to daily life matters\n✍️ Writing Assistance – Draft essays, emails and reports, or polish your written text\n💻 Coding Support – Write and debug code, and explain technical concepts\n🔍 Analysis & Summarization – Process lengthy documents, extract key points and analyze data\n🌍 Translation – Support mutual translation between multiple languages\n💡 Brainstorming – Help you expand ideas and generate creative inspirations\nYou also accept extremely long context inputs, making you ideal for users who need analysis or summaries of lengthy documents." }, { "role": "user", "content": "Calculate fuel consumption." }, { "role": "system", "tools": [ { "type": "function", "function": { "name": "Calculator", "description": "A calculator that evaluates a single arithmetic expression", "parameters": { "type": "object", "properties": { "expr": { "type": "string", "description": "An arithmetic expression in JavaScript syntax; supports basic arithmetic, exponentiation, logarithms, and trigonometric functions" } }, "required": ["expr"] } } } ] } ] } ``` A few things to know: * A `system` message carrying `tools` has the **same standing as ordinary input messages**: the tools become visible to the model starting from the position where the message appears in the `messages` list; * Dynamically loaded tools **coexist** with the global tools declared in the top-level `tools` field, and the model can see both; * A dynamically injected tool declaration must be a **complete** tool definition; you cannot pass only a tool name or reference a tool already declared globally. ```bash theme={null} $ curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k3", "messages": [ { "role": "system", "content": "You are Kimi, an AI assistant developed by Moonshot AI.\nYou are capable of a wide range of tasks, including:\n📝 Q&A & Explanation – Answer all kinds of questions, ranging from scientific knowledge to daily life matters\n✍️ Writing Assistance – Draft essays, emails and reports, or polish your written text\n💻 Coding Support – Write and debug code, and explain technical concepts\n🔍 Analysis & Summarization – Process lengthy documents, extract key points and analyze data\n🌍 Translation – Support mutual translation between multiple languages\n💡 Brainstorming – Help you expand ideas and generate creative inspirations\nYou also accept extremely long context inputs, making you ideal for users who need analysis or summaries of lengthy documents." }, { "role": "user", "content": "Help me compute 23 * 47." }, { "role": "system", "tools": [ { "type": "function", "function": { "name": "Calculator", "description": "A calculator that evaluates a single arithmetic expression", "parameters": { "type": "object", "properties": { "expr": { "type": "string", "description": "An arithmetic expression in JavaScript syntax; supports basic arithmetic, exponentiation, logarithms, and trigonometric functions" } }, "required": ["expr"] } } } ] } ] }' ``` ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "system", "content": "You are Kimi, an AI assistant developed by Moonshot AI.\nYou are capable of a wide range of tasks, including:\n📝 Q&A & Explanation – Answer all kinds of questions, ranging from scientific knowledge to daily life matters\n✍️ Writing Assistance – Draft essays, emails and reports, or polish your written text\n💻 Coding Support – Write and debug code, and explain technical concepts\n🔍 Analysis & Summarization – Process lengthy documents, extract key points and analyze data\n🌍 Translation – Support mutual translation between multiple languages\n💡 Brainstorming – Help you expand ideas and generate creative inspirations\nYou also accept extremely long context inputs, making you ideal for users who need analysis or summaries of lengthy documents."}, {"role": "user", "content": "Help me compute 23 * 47."}, # Dynamically load tools: insert a system message carrying a tools field { "role": "system", "tools": [ { "type": "function", "function": { "name": "Calculator", "description": "A calculator that evaluates a single arithmetic expression", "parameters": { "type": "object", "properties": { "expr": { "type": "string", "description": "An arithmetic expression in JavaScript syntax; supports basic arithmetic, exponentiation, logarithms, and trigonometric functions", } }, "required": ["expr"], }, }, } ], }, ], ) print(completion.choices[0].message.tool_calls) ``` ## Implementing tool search with dynamic loading There is no dedicated tool-search API. If you have a large tool inventory, you can implement tool search yourself by combining a **custom search tool with dynamically loaded tools**: 1. Declare only a single `search_tools` function in the top-level `tools` field, implemented by your backend, which returns matching tool names and summaries for a given keyword; 2. In the system prompt, advertise the searchable keywords (e.g. a tool catalog or domain tags) so the model knows to call `search_tools` first when it needs a tool; 3. Based on what `search_tools` returns, your application inserts the **full declarations** of the matching tools into `messages` via a `system` message carrying a `tools` field; 4. The model can then call these newly loaded tools in subsequent generations. No matter how large the total tool inventory is, each request then only carries a handful of tool declarations, keeping both the context window and the model's selection pressure under control. ## Impact on context caching Dynamically loaded tools can be combined with [Context Caching](/docs/guide/use-context-caching-feature-of-kimi-api). Context caching works by prefix matching: only the leading portion of the current request that is identical to a previous request can hit the cache, and any change within the prefix invalidates the cache from that point onward. How you inject tool declarations therefore directly determines your cache hit rate. Follow these principles to keep a high hit rate while loading tools on demand: * **Append, never insert**: always append new tool declarations to the end of `messages`. The existing prefix stays unchanged and the established cache is unaffected. Inserting or modifying any message in the middle of the conversation (including an already injected tool declaration) invalidates the cache from that point onward; * **Keep injected declarations**: dynamic tool declarations apply per request and are not retained by the server. We recommend carrying previously loaded declarations unchanged in subsequent requests, as this keeps the tools available and preserves a stable prefix for consistent cache hits; you may, however, adjust this behavior according to your business requirements. If a declaration is omitted, it ceases to apply, and the model cannot invoke that tool unless it is declared elsewhere. In addition, because `messages` has changed, the prefix following that point may no longer hit the cache; * **Pin core tools at the top level and leave them unchanged**: declare the tools you need every turn as global tools in the top-level `tools` field, and keep them unchanged afterwards. Top-level global tool declarations do not affect cache hits, so keeping them stable preserves the effectiveness of the prefix cache. Use dynamic injection only for on-demand tools. | Operation | Effect on the prefix cache | | :------------------------------------------------------------------------------------------------ | :------------------------------------------------- | | Appending a tool declaration at the end of `messages` | Existing prefix cache unaffected | | Carrying previously injected declarations unchanged | Prefix remains stable, sustaining cache hits | | Deleting or modifying a message mid-conversation, or inserting a new declaration mid-conversation | Cache after the point of change may be invalidated | | Declaring global tools in the top-level `tools` field | Cache hits unaffected | Note the threshold for caching to take effect: a new request can hit the prefix cache only when the previous request's prompt tokens exceed 256; below 256 tokens the request is not cached and is discarded. See [Context Caching](/docs/guide/use-context-caching-feature-of-kimi-api) for details. ## Notes * Dynamic tool declarations use **exactly the same format** as global `tools` declarations, so you maintain a single schema and migration stays cheap; * A `system` message carrying `tools` also consumes context length, so only inject the tools the current conversation genuinely needs; * Dynamically loaded tools are currently supported only on `kimi-k3`; on other models (e.g. `kimi-k2.6`) the request fails with a `tokenization failed` error; * A `system` message carrying `tools` must not also carry a `content` field, otherwise the request fails with a 400 error (`cannot be used with content`); with the OpenAI SDK you can pass the `tools` field through directly in `messages`, with no `extra_body` needed. ## Related reading * [Kimi K3 API Tool Calling Best Practices](/docs/guide/kimi-k3-tool-calling-best-practice): combined practices for dynamic loading, tool\_choice, and reasoning effort * [Tool Choice](/docs/guide/use-tool-choice): constrain the model's tool-calling behavior with `tool_choice` * [Use Kimi API for Tool Calls](/docs/guide/use-kimi-api-to-complete-tool-calls): the complete tool-calling workflow and examples * [Model Parameter Reference](/docs/api/models-overview): per-model support for parameters such as `tool_choice` # Use Kimi API's JSON Mode Source: https://platform.kimi.ai/docs/guide/use-json-mode-feature-of-kimi-api Enable Kimi API JSON Mode with `response_format` and safely prompt for, receive, and parse structured JSON output. JSON Mode makes the Kimi large language model output a valid, correctly parsable JSON document. When you need structured output — for example, summarizing an article into structured data like this — enable it with the `response_format` parameter: ```json theme={null} { "title": "Article Title", "author": "Article Author", "publish_time": "Publication Time", "summary": "Article Summary" } ``` ## Enable JSON Mode with response\_format If you only tell the Kimi large language model in the prompt: "Please output content in JSON format," the model can understand your request and generate a JSON document as required. However, the generated content often has some flaws: for instance, in addition to the JSON document, Kimi might output extra text to explain the JSON document — ```text theme={null} Here is the JSON document you requested { "title": "Article Title", "author": "Article Author", "publish_time": "Publication Time", "summary": "Article Summary" } ``` —or the JSON document might be malformed and cannot be parsed correctly (note the comma at the end of the `summary` field on the last line): ```text theme={null} { "title": "Article Title", "author": "Article Author", "publish_time": "Publication Time", "summary": "Article Summary", } ``` The `response_format` parameter constrains the output format. Its default value is `{"type": "text"}`, which means ordinary text content with no formatting constraints. Set `response_format` to `{"type": "json_object"}` to enable JSON Mode, and the Kimi large language model will output a valid, correctly parsable JSON document as required. Using JSON Mode takes three steps: 1. Define the output JSON format in the system or user prompt, including specific field names and field types; **the best practice is to provide a concrete output example and explain the meaning of each field**; 2. Set the `response_format` parameter to `{"type": "json_object"}`; 3. Parse the `content` in the message returned by the Kimi large language model; `message.content` is a valid JSON Object serialized as a string. ## Full example: a customer-service bot with mixed message types Imagine a WeChat intelligent robot customer service (referred to as intelligent customer service): it uses the Kimi large language model to answer customer questions, and can reply not only with text messages but also with images, link cards, voice messages, and other types of messages, mixing different types of messages in a single response. For example, for customer product inquiries, it provides a text reply, a product image, and finally a purchase link (in the form of a link card). The following code demonstrates how to use JSON Mode in this scenario to make the model output replies in a fixed structure, and how to parse each type of message in the returned content: The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ```python theme={null} import os import json from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url="https://api.moonshot.ai/v1", ) system_prompt = """ You are the intelligent customer service of Moonshot AI (Kimi), responsible for answering various user questions. Please refer to the document content to reply to user questions. Your reply can be text, images, links, and you can include text, images, and links in a single response. Please output your reply in the following JSON format: { "text": "Text information", "image": "Image URL", "url": "Link URL" } Note: Please place the text information in the `text` field, put the image in the `image` field in the form of a link starting with `oss://`, and place the regular link in the `url` field. """ completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "system", "content": "You are Kimi, an artificial intelligence assistant provided by Moonshot AI, excelling in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You will reject any questions involving terrorism, racism, pornography, and violence. Moonshot AI is a proper noun and should not be translated into other languages."}, {"role": "system", "content": system_prompt}, # <-- Submit the system prompt with the output format to Kimi {"role": "user", "content": "Hello, my name is Li Lei, what is 1+1?"} ], response_format={"type": "json_object"}, # <-- Use the response_format parameter to specify the output format as json_object ) # Since we have set JSON Mode, the message.content returned by the Kimi large language model is a serialized JSON Object string. # We use json.loads to parse its content and deserialize it into a Python dictionary. content = json.loads(completion.choices[0].message.content) # Parse text content if "text" in content: # For demonstration purposes, we print the content; # In real business logic, you may need to call the text message sending interface to send the generated text to the user. print("text:", content["text"]) # Parse image content if "image" in content: # For demonstration purposes, we print the content; # In real business logic, you may need to first parse the image URL, download the image, and then call the image message sending # interface to send the image to the user. print("image:", content["image"]) # Parse link if "url" in content: # For demonstration purposes, we print the content; # In real business logic, you may need to call the link card sending interface to send the link to the user in the form of a card. print("url:", content["url"]) ``` ```js theme={null} const OpenAI = require("openai") const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }) system_prompt = ` You are the intelligent customer service of Moonshot AI (Kimi), responsible for answering various user questions. Please refer to the document content to reply to user questions. Your reply can be text, images, links, and you can include text, images, and links in a single response. " " Please output your reply in the following JSON format: { "text": "Text information", "image": "Image URL", "url": "Link URL" } " Note: Please place the text information in the 'text' field, put the image in the 'image' field in the form of a link starting with oss://, and place the regular link in the 'url' field. ` async function main() { const completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ {role: "system", content: "You are Kimi, an artificial intelligence assistant provided by Moonshot AI, excelling in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You will reject any questions involving terrorism, racism, pornography, and violence. Moonshot AI is a proper noun and should not be translated into other languages."}, {role: "system", content: system_prompt}, // <-- Submit the system prompt with the output format to Kimi {role: "user", content: "Hello, my name is Li Lei, what is 1+1?"} ], response_format: {type: "json_object"}, // <-- Use the response_format parameter to specify the output format as json_object }) // Since we have set JSON Mode, the message.content returned by the Kimi large language model is a serialized JSON Object string. // We use JSON.parse to parse its content and deserialize it into a JavaScript object. content = JSON.parse(completion.choices[0].message.content) // Parse text content if (content.text) { // For demonstration purposes, we print the content; // In real business logic, you may need to call the text message sending interface to send the generated text to the user. console.log("text:", content.text) } // Parse image content if (content.image) { // For demonstration purposes, we print the content; // In real business logic, you may need to first parse the image URL, download the image, and then call the image message sending // interface to send the image to the user. console.log("image:", content.image) } // Parse link if (content.url) { // For demonstration purposes, we print the content; // In real business logic, you may need to call the link card sending interface to send the link to the user in the form of a card. console.log("url", content.url) } } main() ``` ## Troubleshoot truncated JSON output If you have correctly set the `response_format` parameter and specified the format of the JSON document in the prompt, but the JSON document you receive is incomplete or truncated and cannot be parsed correctly, check whether the `finish_reason` field in the return value is `length`. A smaller `max_tokens` value will cause the model's output to be truncated, and this rule also applies when using JSON Mode. We recommend estimating the size of the output JSON document and setting a reasonable `max_tokens` value, so that you can correctly parse the JSON document returned by the Kimi large language model. For a more detailed explanation of incomplete or truncated output from the Kimi large language model, see [Troubleshooting](/docs/guide/troubleshooting). ## Notes * The Kimi large language model only generates JSON Object type JSON documents; do not prompt it to generate JSON Array or other types of JSON documents. * If you do not correctly inform the Kimi large language model of the required JSON Object format, it will generate unexpected results. # Use Kimi API for File-Based Q&A Source: https://platform.kimi.ai/docs/guide/use-kimi-api-for-file-based-qa Upload files, extract their contents, and add them to Kimi API conversations for single-file or multi-file question answering. The Kimi intelligent assistant can upload files and answer questions based on those files. The Kimi API offers the same functionality. Below, we'll walk through a practical example of how to upload files and ask questions using the Kimi API: The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ```python theme={null} import os from pathlib import Path from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url="https://api.moonshot.ai/v1", ) # 'moonshot.pdf' is an example file. We support text and image files. For image files, we provide OCR capabilities. # To upload a file, you can use the file upload API from the openai library. Create a file object using Path from the standard library pathlib and pass it to the file parameter. Set the purpose parameter to 'file-extract'. The file upload interface also supports other purpose values such as 'image' and 'video' (for native model understanding). file_object = client.files.create(file=Path("moonshot.pdf"), purpose="file-extract") # Get the result # file_content = client.files.retrieve_content(file_id=file_object.id) # Note: The retrieve_content API in some older examples is marked as deprecated in the latest version. You can use the following line instead (if you're using an older SDK version, you can continue using retrieve_content). file_content = client.files.content(file_id=file_object.id).text # Include the file content in the request as a system prompt messages = [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You excel in Chinese and English conversations. You provide users with safe, helpful, and accurate answers while rejecting any queries related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated.", }, { "role": "system", "content": file_content, # <-- Here, we place the extracted file content (note that it's the content, not the file ID) in the request }, {"role": "user", "content": "Please give a brief introduction to the content of moonshot.pdf"}, ] # Then call the chat-completion API to get Kimi's response completion = client.chat.completions.create( model="kimi-k3", messages=messages ) print(completion.choices[0].message) ``` ```js theme={null} const OpenAI = require("openai"); const fs = require("fs") const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }); async function main() { // 'moonshot.pdf' is an example file. We support pdf, doc, and image formats, providing OCR capabilities for images and pdf files. let file_object = await client.files.create({ file: fs.createReadStream("moonshot.pdf"), purpose: "file-extract" }) // Get the result // file_content = client.files.retrieve_content(file_id=file_object.id) // Note: The retrieve_content API in some older examples is marked as deprecated in the latest version. You can use the following line instead (if you're using an older SDK version, you can continue using retrieve_content). let file_content = await (await client.files.content(file_object.id)).text() // Include it in the request let messages = [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You excel in Chinese and English conversations. You provide users with safe, helpful, and accurate answers while rejecting any queries related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated.", }, { "role": "system", "content": file_content, }, {"role": "user", "content": "Please give a brief introduction to the content of moonshot.pdf"}, ] const completion = await client.chat.completions.create({ model: "kimi-k3", messages: messages }); console.log(completion.choices[0].message.content); } main(); ``` Let's review the basic steps and considerations for file-based Q\&A: 1. Upload the file to the Kimi server using the `/v1/files` interface or the `files.create` API in the SDK; 2. Retrieve the file content using the `/v1/files/{file_id}` interface or the `files.content` API in the SDK. The retrieved content is already formatted in a way that our recommended model can easily understand; 3. Place the extracted (and formatted) file content (not the file `id`) in the messages list as a system prompt; 4. Start asking questions about the file content; **Note again: Place the file content in the prompt, not the file `id`.** ## Q\&A on Multiple Files If you want to ask questions based on multiple files, it's quite simple. **Just place each file in a separate system prompt.** Here's how you can do it in code: ```python theme={null} from typing import * import os import json from pathlib import Path from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url="https://api.moonshot.ai/v1", ) def upload_files(files: List[str]) -> List[Dict[str, Any]]: """ upload_files uploads all the provided files (paths) via the '/v1/files' interface and generates file messages from the extracted content. Each file becomes an independent message with a role of 'system', which the Kimi large language model can correctly identify. :param files: A list of file paths to be uploaded. The paths can be absolute or relative, and should be passed as strings. :return: A list of messages containing the file content. Add these messages to the Context, i.e., the messages parameter when calling the `/v1/chat/completions` interface. """ messages = [] # For each file path, we upload the file, extract its content, and generate a message with a role of 'system', which is then added to the final messages list. for file in files: file_object = client.files.create(file=Path(file), purpose="file-extract") file_content = client.files.content(file_id=file_object.id).text messages.append({ "role": "system", "content": file_content, }) return messages def main(): file_messages = upload_files(files=["upload_files.py"]) messages = [ # We use the * syntax to unpack the file_messages, making them the first N messages in the messages list. *file_messages, { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You excel in Chinese and English conversations. You provide users with safe, helpful, and accurate answers while rejecting any queries related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated.", }, { "role": "user", "content": "Summarize the content of these files.", }, ] print(json.dumps(messages, indent=2, ensure_ascii=False)) completion = client.chat.completions.create( model="kimi-k3", messages=messages, ) print(completion.choices[0].message.content) if __name__ == '__main__': main() ``` ```js theme={null} const fs = require('fs'); const path = require('path'); const axios = require('axios'); const OpenAI = require('openai'); const client = new OpenAI({ baseURL: "https://api.moonshot.ai/v1", // Set the MOONSHOT_API_KEY environment variable before running this example. apiKey: process.env.MOONSHOT_API_KEY, }) async function upload_files(files){ /* The upload_files function uploads all the provided file paths through the file upload interface '/v1/files', and generates a list of messages from the uploaded file contents. Each file will be a separate message, and all these messages will have the role of 'system'. The Kimi large language model will correctly recognize the file content within these system messages. We recommend placing the messages returned by upload_files at the head of the messages list. */ let messages = [] // For each file path, we upload the file, extract its content, and then generate a message with the // role of 'system', which is added to the final list of messages to be returned. for (const file of files) { const file_object = await client.files.create({file: fs.createReadStream(path.resolve(file)), purpose: "file-extract"}) let file_content = await (await client.files.content(file_object.id)).text() messages.push({ role: "system", content: file_content, }) } return messages } async function main() { const fileMessages = await upload_files( ["upload_files.py"] ) const messages = [ // We use the ... syntax to destructure the file_messages, making them the first N messages in // the messages list. ...fileMessages, { role: "system", content: "You are Kimi, an AI assistant provided by Moonshot AI, and you are particularly " + "skilled in Chinese and English conversations. You provide users with safe, helpful, " + "and accurate answers. At the same time, you refuse to answer any questions " + "involving terrorism, racism, pornography, or violence. Moonshot AI is a proper " + "noun and should not be translated into other languages.", }, { role: "user", content: "Summarize the content of these files.", } ] console.log(JSON.stringify(messages, null, 2)) const completion = await client.chat.completions.create({ model: "kimi-k3", messages: messages, }) console.log(completion.choices[0].message.content) } main() ``` ## Best Practices for File Management In general, the file upload and extraction features are designed to convert files of various formats into a format that our recommended model can easily understand. After completing the file upload and extraction steps, the extracted content can be stored locally. In the next file-based Q\&A request, there is no need to upload and extract the files again. Since we have limited the number of files a single user can upload (up to 1000 files per user), we suggest that you regularly clean up the uploaded files after the extraction process is complete. You can periodically run the following code to clean up the uploaded files: ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url="https://api.moonshot.ai/v1", ) file_list = client.files.list() for file in file_list.data: client.files.delete(file_id=file.id) ``` ```js theme={null} const OpenAI = require("openai") const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }) async function main() { const file_list = await client.files.list() for (file of file_list.data) { await client.files.delete(file_id=file.id) } } main() ``` In the code above, we first list all the file details using the `files.list` API and then delete each file using the `files.delete` API. Regularly performing this operation ensures that file storage space is released, allowing subsequent file uploads and extractions to be successful. # Use Kimi API for Tool Calls Source: https://platform.kimi.ai/docs/guide/use-kimi-api-to-complete-tool-calls Define, register, and execute Kimi API `tool_calls`, return tool results, and handle tool calls in streaming responses. Tool calls (`tool_calls`) let the Kimi large language model go beyond "talking" to "doing": the model decides whether to call a tool based on the conversation context, generates the call arguments in JSON, and your application executes the tool and returns the result so the model can produce the final reply. With `tool_calls`, the Kimi large language model can help you search the internet, query databases, and even control smart home devices. This page walks through the full flow — defining, registering, and executing tools — with an internet-search example, plus notes for streaming and other scenarios. ## The Complete Flow of a Tool Call A tool call involves the following steps: 1. Define the tool using JSON Schema format; 2. Submit the defined tool to the Kimi large language model via the `tools` parameter. You can submit multiple tools at once; 3. The Kimi large language model will decide which tool(s) to use based on the context of the current conversation. It can also choose not to use any tools; 4. The Kimi large language model will output the parameters and information needed to call the tool in JSON format; 5. Use the parameters output by the Kimi large language model to execute the corresponding tool and submit the results back to the Kimi large language model; 6. The Kimi large language model will respond to the user based on the results of the tool execution; If your application needs a large tool inventory (dozens or hundreds of tools), use [Dynamically Loaded Tools](/docs/guide/use-dynamic-tool-loading) to inject tool definitions on demand instead of submitting them all at once — this significantly reduces token usage and improves tool-selection accuracy. ## Give the Model Internet Access with Tool Calls The knowledge of the Kimi large language model comes from its training data, so it cannot answer time-sensitive questions from what it already knows. The example below uses two tools — a "search engine" and a "web browser" — to show how the model can search for the latest information and answer based on it. ### Define Tools with JSON Schema When people look up information online, they usually open a search engine (such as Baidu or Bing), browse the search results, and then open one or more result pages to get the knowledge they need. Abstracting these two actions into tools gives us a "search engine" and a "web browser" — described in JSON Schema and submitted to the Kimi large language model, so it can search and browse the web just like humans do. Tool definitions are written in JSON Schema format: > [JSON Schema](https://json-schema.org/) is a vocabulary that you can use to annotate and validate JSON documents. > > [JSON Schema](https://json-schema.org/) is a JSON document used to describe the format of JSON data. We define the following JSON Schema: ```json theme={null} { "type": "object", "properties": { "name": { "type": "string" } } } ``` This JSON Schema defines a JSON Object that contains a field named `name`, and the type of this field is `string`, for example: ```json theme={null} { "name": "Hei" } ``` By describing our tool definitions using JSON Schema, we can make it clearer and more intuitive for the Kimi large language model to understand what parameters our tools require, as well as the type and description of each parameter. Now let's define the "search engine" and "web browser" tools mentioned earlier: ```python theme={null} tools = [ { "type": "function", # The agreed-upon field type, currently supports function as a value "function": { # When type is function, use the function field to define the specific function content "name": "search", # The name of the function. Please use English letters, numbers, hyphens, and underscores as the function name "description": """ Search for content on the internet using a search engine. When your knowledge cannot answer the user's question, or when the user requests an online search, call this tool. Extract the content the user wants to search for from the conversation and use it as the value of the query parameter. The search results include the website title, address (URL), and description. """, # A description of the function, detailing its specific role and usage scenarios, to help the Kimi large language model correctly select which functions to use "parameters": { # Use the parameters field to define the parameters the function accepts "type": "object", # Always use type: object to make the Kimi large language model generate a JSON Object parameter "required": ["query"], # Use the required field to tell the Kimi large language model which parameters are mandatory "properties": { # The properties field contains the specific parameter definitions; you can define multiple parameters "query": { # Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", # Use type to define the parameter type "description": """ The content the user wants to search for, extracted from the user's question or conversation context. """ # Use description to describe the parameter so that the Kimi large language model can better generate the parameter } } } } }, { "type": "function", # The agreed-upon field type, currently supports function as a value "function": { # When type is function, use the function field to define the specific function content "name": "crawl", # The name of the function. Please use English letters, numbers, hyphens, and underscores as the function name "description": """ Retrieve web page content based on the website address (URL). """, # A description of the function, detailing its specific role and usage scenarios, to help the Kimi large language model correctly select which functions to use "parameters": { # Use the parameters field to define the parameters the function accepts "type": "object", # Always use type: object to make the Kimi large language model generate a JSON Object parameter "required": ["url"], # Use the required field to tell the Kimi large language model which parameters are mandatory "properties": { # The properties field contains the specific parameter definitions; you can define multiple parameters "url": { # Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", # Use type to define the parameter type "description": """ The website address (URL) from which to retrieve content, usually obtained from search results. """ # Use description to describe the parameter so that the Kimi large language model can better generate the parameter } } } } } ] ``` ```js theme={null} const tools = [ { "type": "function", // The field "type" is a convention, currently supporting "function" as its value "function": { // When "type" is "function", use the "function" field to define the specific function content "name": "search", // The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name "description": ""/* Search for content on the internet using a search engine. When your knowledge cannot answer the user's question, or when the user requests an online search, call this tool. Extract the content the user wants to search from the conversation as the value of the query parameter. The search results include the website title, website address (URL), and website description. */, // Description of the function, write the specific function and usage scenarios here so that the Kimi large language model can correctly choose which functions to use "parameters": { // Use the "parameters" field to define the parameters accepted by the function "type": "object", // Always use "type": "object" to make the Kimi large language model generate a JSON Object parameter "required": ["query"], // Use the "required" field to tell the Kimi large language model which parameters are required "properties": { // The specific parameter definitions are in "properties", you can define multiple parameters "query": { // Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", // Use "type" to define the parameter type "description": ""/* The content the user wants to search for, extract it from the user's question or chat context. */ // Use "description" to describe the parameter so that the Kimi large language model can better generate the parameter } } } } }, { "type": "function", // The field "type" is a convention, currently supporting "function" as its value "function": { // When "type" is "function", use the "function" field to define the specific function content "name": "crawl", // The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name "description": ""/* Get the content of a webpage based on the website address (URL). */, // Description of the function, write the specific function and usage scenarios here so that the Kimi large language model can correctly choose which functions to use "parameters": { // Use the "parameters" field to define the parameters accepted by the function "type": "object", // Always use "type": "object" to make the Kimi large language model generate a JSON Object parameter "required": ["url"], // Use the "required" field to tell the Kimi large language model which parameters are required "properties": { // The specific parameter definitions are in "properties", you can define multiple parameters "url": { // Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", // Use "type" to define the parameter type "description": ""/* The website address (URL) of the content to be obtained, which can usually be obtained from the search results. */ // Use "description" to describe the parameter so that the Kimi large language model can better generate the parameter } } } } } ] ``` When defining tools using JSON Schema, we use the following fixed format: ```json theme={null} { "type": "function", "function": { "name": "NAME", "description": "DESCRIPTION", "parameters": { "type": "object", "properties": { } } } } ``` Here, `name`, `description`, and `parameters.properties` are defined by the tool provider. The `description` explains the specific function and when to use the tool, while `parameters` outlines the specific parameters needed to successfully call the tool, including parameter types and descriptions. **Ultimately, the Kimi large language model will generate a JSON Object that meets the defined requirements as the parameters (arguments) for the tool call based on the JSON Schema.** ### Register the Tools with the Model Submit the `search` tool to the Kimi large language model and see if it can call the tool correctly: The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url="https://api.moonshot.ai/v1", ) tools = [ { "type": "function", # The field "type" is a convention, currently supporting "function" as its value "function": { # When "type" is "function", use the "function" field to define the specific function content "name": "search", # The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name "description": """ Search for content on the internet using a search engine. When your knowledge cannot answer the user's question, or when the user requests an online search, call this tool. Extract the content the user wants to search from the conversation as the value of the query parameter. The search results include the website title, website address (URL), and website description. """, # Description of the function, write the specific function and usage scenarios here so that the Kimi large language model can correctly choose which functions to use "parameters": { # Use the "parameters" field to define the parameters accepted by the function "type": "object", # Always use "type": "object" to make the Kimi large language model generate a JSON Object parameter "required": ["query"], # Use the "required" field to tell the Kimi large language model which parameters are required "properties": { # The specific parameter definitions are in "properties", you can define multiple parameters "query": { # Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", # Use "type" to define the parameter type "description": """ The content the user wants to search for, extract it from the user's question or chat context. """ # Use "description" to describe the parameter so that the Kimi large language model can better generate the parameter } } } } }, # { # "type": "function", # The field "type" is a convention, currently supporting "function" as its value # "function": { # When "type" is "function", use the "function" field to define the specific function content # "name": "crawl", # The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name # "description": """ # Get the content of a webpage based on the website address (URL). # """, // Description of the function, write the specific function and usage scenarios here so that the Kimi large language model can correctly choose which functions to use # "parameters": { // Use the "parameters" field to define the parameters accepted by the function # "type": "object", // Always use "type": "object" to make the Kimi large language model generate a JSON Object parameter # "required": ["url"], // Use the "required" field to tell the Kimi large language model which parameters are required # "properties": { // The specific parameter definitions are in "properties", you can define multiple parameters # "url": { // Here, the key is the parameter name, and the value is the specific definition of the parameter # "type": "string", // Use "type" to define the parameter type # "description": """ # The website address (URL) of the content to be obtained, which can usually be obtained from the search results. # """ // Use "description" to describe the parameter so that the Kimi large language model can better generate the parameter # } # } # } # } # } ] completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You refuse to answer any questions related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated."}, {"role": "user", "content": "Please search the internet for 'Context Caching' and tell me what it is."} # In the question, we ask Kimi large language model to search online ], tools=tools, # <-- We pass the defined tools to Kimi large language model via the tools parameter ) print(completion.choices[0].model_dump_json(indent=4)) ``` ```js theme={null} const OpenAI = require("openai") const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }) const tools = [ { "type": "function", // The agreed-upon field type, currently supports function as a value "function": { // When type is function, use the function field to define the specific function content "name": "search", // The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name "description": ""/* Search for content on the internet using a search engine. When your knowledge cannot answer the user's question, or the user requests you to search online, call this tool. Extract the content the user wants to search from the conversation as the value of the query parameter. The search results include the website title, website address (URL), and website description. */, // Description of the function, write the specific function's role and usage scenario here to help Kimi large language model correctly choose which functions to use "parameters": { // Use the parameters field to define the parameters the function accepts "type": "object", // Fixedly use type: object to make Kimi large language model generate a JSON Object parameter "required": ["query"], // Use the required field to tell Kimi large language model which parameters are required "properties": { // The specific parameter definitions are in properties, you can define multiple parameters "query": { // Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", // Use type to define the parameter type "description": ""/* The content the user wants to search for, extracted from the user's question or chat context. */ // Use description to describe the parameter so that Kimi large language model can better generate the parameter } } } } }, // { // "type": "function", // The agreed-upon field type, currently supports function as a value // "function": { // When type is function, use the function field to define the specific function content // "name": "crawl", // The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name // "description": """ // Get the webpage content based on the website address (URL). // """, // Description of the function, write the specific function's role and usage scenario here to help Kimi large language model correctly choose which functions to use // "parameters": { // Use the parameters field to define the parameters the function accepts // "type": "object", // Fixedly use type: object to make Kimi large language model generate a JSON Object parameter // "required": ["url"], // Use the required field to tell Kimi large language model which parameters are required // "properties": { // The specific parameter definitions are in properties, you can define multiple parameters // "url": { // Here, the key is the parameter name, and the value is the specific definition of the parameter // "type": "string", // Use type to define the parameter type // "description": """ // The website address (URL) whose content needs to be obtained, usually the website address can be obtained from the search results. // """ // Use description to describe the parameter so that Kimi large language model can better generate the parameter // } // } // } // } // } ] async function main() { const completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ {role: "system", content: "You are Kimi, an AI assistant provided by Moonshot AI. You are proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You refuse to answer any questions related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated."}, {role: "user", content: "Please search the internet for 'Context Caching' and tell me what it is."} // In the question, we ask Kimi large language model to search online ], tools: tools, // <-- We pass the defined tools to Kimi large language model via the tools parameter }) console.log(JSON.stringify(completion.choices[0], null, 4)) } main() ``` When the above code runs successfully, we get the response from Kimi large language model: ```json theme={null} { "finish_reason": "tool_calls", "message": { "content": "", "role": "assistant", "tool_calls": [ { "id": "search:0", "function": { "arguments": "{\n \"query\": \"Context Caching\"\n}", "name": "search" }, "type": "function" } ] } } ``` Notice that in this response, the value of `finish_reason` is `tool_calls`, which means that the response is not the answer from Kimi large language model, but rather the tool that Kimi large language model has chosen to execute. You can determine whether the current response from Kimi large language model is a tool call `tool_calls` by checking the value of `finish_reason`. At this point the `content` field in `message` is empty, because the model is executing `tool_calls` and has not yet generated a reply for the user. The newly added `tool_calls` field is a list containing all the tool call information for this turn — which shows that **the model can choose to call multiple tools at once, which can be different tools or the same tool with different parameters**. Each element in `tool_calls` represents one tool call: the Kimi large language model generates a unique `id` for each call, uses `function.name` to indicate the name of the tool function, and places the call parameters in `function.arguments` (`arguments` is a valid serialized JSON Object; additionally, the `type` parameter is currently a fixed value `function`). Next, use the tool call parameters generated by the Kimi large language model to execute the corresponding tools. ### Execute the Tools and Return the Results The Kimi large language model does not execute tools for you — once you receive the parameters it generates, your application must execute them. Why can't the model execute tools itself? Imagine a typical scenario: **you provide users with a smart robot based on the Kimi large language model. In this scenario, there are three roles: the user, the robot, and the Kimi large language model. The user asks the robot a question, the robot calls the Kimi large language model API, and returns the API result to the user. When using `tool_calls`, the user asks the robot a question, the robot calls the Kimi API with `tools`, the Kimi large language model returns the `tool_calls` parameters, the robot executes the `tool_calls`, submits the results back to the Kimi API, the Kimi large language model generates the message to be returned to the user (`finish_reason=stop`), and only then does the robot return the message to the user.** The entire `tool_calls` process is transparent and implicit to the user: users never directly "see" the tool calls, only the final reply presented by the robot. The full example below executes the `tool_calls` returned by the Kimi large language model from the perspective of the "robot", demonstrating the tool-execution loop: ```python theme={null} from typing import * import json import httpx import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url="https://api.moonshot.ai/v1", ) tools = [ { "type": "function", # The field type is agreed upon, and currently supports function as a value "function": { # When type is function, use the function field to define the specific function content "name": "search", # The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name "description": """ Search for content on the internet using a search engine. When your knowledge cannot answer the user's question, or the user requests you to perform an online search, call this tool. Extract the content the user wants to search from the conversation as the value of the query parameter. The search results include the title of the website, the website address (URL), and a brief introduction to the website. """, # Introduction to the function, write the specific function here, as well as the usage scenario, so that the Kimi large language model can correctly choose which functions to use "parameters": { # Use the parameters field to define the parameters accepted by the function "type": "object", # Fixed use type: object to make the Kimi large language model generate a JSON Object parameter "required": ["query"], # Use the required field to tell the Kimi large language model which parameters are required "properties": { # The specific parameter definitions are in properties, and you can define multiple parameters "query": { # Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", # Use type to define the parameter type "description": """ The content the user wants to search for, extracted from the user's question or chat context. """ # Use description to describe the parameter so that the Kimi large language model can better generate the parameter } } } } }, { "type": "function", # The field type is agreed upon, and currently supports function as a value "function": { # When type is function, use the function field to define the specific function content "name": "crawl", # The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name "description": """ Get the content of a webpage based on the website address (URL). """, # Introduction to the function, write the specific function here, as well as the usage scenario, so that the Kimi large language model can correctly choose which functions to use "parameters": { # Use the parameters field to define the parameters accepted by the function "type": "object", # Fixed use type: object to make the Kimi large language model generate a JSON Object parameter "required": ["url"], # Use the required field to tell the Kimi large language model which parameters are required "properties": { # The specific parameter definitions are in properties, and you can define multiple parameters "url": { # Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", # Use type to define the parameter type "description": """ The website address (URL) of the content to be obtained, which can usually be obtained from the search results. """ # Use description to describe the parameter so that the Kimi large language model can better generate the parameter } } } } } ] def search_impl(query: str) -> List[Dict[str, Any]]: """ search_impl uses a search engine to search for query. Most mainstream search engines (such as Bing) provide API calls. You can choose your preferred search engine API and place the website title, link, and brief introduction information from the return results in a dict to return. This is just a simple example, and you may need to write some authentication, validation, and parsing code. """ r = httpx.get("https://your.search.api", params={"query": query}) return r.json() def search(arguments: Dict[str, Any]) -> Any: query = arguments["query"] result = search_impl(query) return {"result": result} def crawl_impl(url: str) -> str: """ crawl_url gets the content of a webpage based on the url. This is just a simple example. In actual web scraping, you may need to write more code to handle complex situations, such as asynchronously loaded data; and after obtaining the webpage content, you can clean the webpage content according to your needs, such as retaining only the text or removing unnecessary content (such as advertisements). """ r = httpx.get(url) return r.text def crawl(arguments: dict) -> str: url = arguments["url"] content = crawl_impl(url) return {"content": content} # Map each tool name and its corresponding function through tool_map so that when the Kimi large language model returns tool_calls, we can quickly find the function to execute tool_map = { "search": search, "crawl": crawl, } messages = [ {"role": "system", "content": "You are Kimi, an artificial intelligence assistant provided by Moonshot AI. You are better at conversing in Chinese and English. You provide users with safe, helpful, and accurate answers. At the same time, you will refuse to answer any questions involving terrorism, racial discrimination, pornography, and violence. Moonshot AI is a proper noun and should not be translated into other languages."}, {"role": "user", "content": "Please search for Context Caching online and tell me what it is."} # Request Kimi large language model to perform an online search in the question ] finish_reason = None # Our basic process is to ask the Kimi large language model questions with the user's question and tools. If the Kimi large language model returns finish_reason: tool_calls, we execute the corresponding tool_calls, # and submit the execution results in the form of a message with role=tool back to the Kimi large language model. The Kimi large language model then generates the next content based on the tool_calls results: # # 1. If the Kimi large language model believes that the current tool call results can answer the user's question, it returns finish_reason: stop, and we exit the loop and print out message.content; # 2. If the Kimi large language model believes that the current tool call results cannot answer the user's question and needs to call the tool again, we continue to execute the next tool_calls in the loop until finish_reason is no longer tool_calls; # # During this process, we only return the result to the user when finish_reason is stop. while finish_reason is None or finish_reason == "tool_calls": completion = client.chat.completions.create( model="kimi-k3", messages=messages, tools=tools, # <-- We submit the defined tools to the Kimi large language model through the tools parameter ) choice = completion.choices[0] finish_reason = choice.finish_reason if finish_reason == "tool_calls": # <-- Determine whether the current return content contains tool_calls messages.append(choice.message) # <-- We add the assistant message returned to us by the Kimi large language model to the context so that the Kimi large language model can understand our request next time for tool_call in choice.message.tool_calls: # <-- tool_calls may be multiple, so we use a loop to execute them one by one tool_call_name = tool_call.function.name tool_call_arguments = json.loads(tool_call.function.arguments) # <-- arguments is a serialized JSON Object, and we need to deserialize it with json.loads tool_function = tool_map[tool_call_name] # <-- Quickly find which function to execute through tool_map tool_result = tool_function(tool_call_arguments) # Construct a message with role=tool using the function execution result to show the result of the tool call to the model; # Note that we need to provide the tool_call_id and name fields in the message so that the Kimi large language model # can correctly match the corresponding tool_call. messages.append({ "role": "tool", "tool_call_id": tool_call.id, "name": tool_call_name, "content": json.dumps(tool_result), # <-- We agree to submit the tool call result to the Kimi large language model in string format, so we use json.dumps to serialize the execution result into a string here }) print(choice.message.content) # <-- Here, we return the reply generated by the model to the user ``` ```js theme={null} const axios = require('axios'); const openai = require('openai'); // You need to install the openai library const client = new openai.OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }); const tools = [ { "type": "function", "function": { "name": "search", "description": "Search for content on the internet using a search engine.\n\nUse this tool when your knowledge can't answer the user's question or when the user asks you to search online. Extract the content the user wants to search for from the conversation and use it as the value for the query parameter.\nThe search results include the website title, URL, and a brief description.", "parameters": { "type": "object", "required": ["query"], "properties": { "query": { "type": "string", "description": "The content the user wants to search for, extracted from the user's question or chat context." } } } } }, { "type": "function", "function": { "name": "crawl", "description": "Retrieve web page content based on a website URL.", "parameters": { "type": "object", "required": ["url"], "properties": { "url": { "type": "string", "description": "The URL of the website whose content you want to retrieve, usually obtained from search results." } } } } } ]; async function searchImpl(query) { const response = await axios.get("https://your.search.api", { params: { query } }); return response.data; } async function search(args) { const query = args.query; const result = await searchImpl(query); return { "result": result }; } async function crawlImpl(url) { const response = await axios.get(url); return response.data; } async function crawl(args) { const url = args.url; const content = await crawlImpl(url); return { "content": content }; } const toolMap = { "search": search, "crawl": crawl, }; const messages = [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You will refuse to answer any questions involving terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated into other languages." }, { "role": "user", "content": "Please search the internet for Context Caching and tell me what it is." } // The user is asking Kimi to search online ]; let finishReason = null; let choice; async function main() { while (finishReason === null || finishReason === "tool_calls") { const completion = await client.chat.completions.create({ model: "kimi-k3", messages: messages, tools: tools, // <-- We pass the defined tools to the Kimi large language model via the tools parameter }); choice = completion.choices[0]; finishReason = choice.finish_reason; if (finishReason === "tool_calls") { // <-- Check if the current response includes tool_calls messages.push(choice.message); // <-- Add the assistant message from Kimi to the context for the next request for (const toolCall of choice.message.tool_calls) { // <-- There might be multiple tool_calls, so we loop through each one const toolCallName = toolCall.function.name; const toolCallArguments = JSON.parse(toolCall.function.arguments); // <-- The arguments are a serialized JSON object, so we need to parse them const toolFunction = toolMap[toolCallName]; // <-- Use tool_map to quickly find which function to execute const toolResult = await toolFunction(toolCallArguments); // Construct a role=tool message with the function execution result to show the model the outcome of the tool call; // Note that we need to provide the tool_call_id and name fields in the message so that Kimi can correctly match the tool_call. messages.push({ "role": "tool", "tool_call_id": toolCall.id, "name": toolCallName, "content": JSON.stringify(toolResult), // <-- We agreed to submit the tool call result as a string, so we serialize it with JSON.stringify }); } } } console.log(choice.message.content); // <-- Finally, we return the model's response to the user } main(); ``` We use a `while` loop to execute the code logic that includes tool calls because the Kimi large language model typically doesn't make just one tool call, especially in the context of online searching. Usually, Kimi will first call the `search` tool to get search results, and then call the `crawl` tool to convert the URLs in the search results into actual web page content. The overall structure of the `messages` is as follows: ```text theme={null} system: prompt # System prompt user: prompt # User's question assistant: tool_call(name=search, arguments={query: query}) # Kimi returns a tool_call (single) tool: search_result(tool_call_id=tool_call.id, name=search) # Submit the tool_call execution result assistant: tool_call_1(name=crawl, arguments={url: url_1}), tool_call_2(name=crawl, arguments={url: url_2}) # Kimi continues to return tool_calls (multiple) tool: crawl_content(tool_call_id=tool_call_1.id, name=crawl) # Submit the execution result of tool_call_1 tool: crawl_content(tool_call_id=tool_call_2.id, name=crawl) # Submit the execution result of tool_call_2 assistant: message_content(finish_reason=stop) # Kimi generates a reply to the user, ending the conversation ``` This completes the entire process of making "online query" tool calls. If you have implemented your own `search` and `crawl` methods, when you ask Kimi to search online, it will call the `search` and `crawl` tools and give you the correct response based on the tool call results. ## Handle tool\_calls in Streaming Output In streaming output mode (`stream`), `tool_calls` work as well, but there are a few extra things to note: * During streaming output, since `finish_reason` will appear in the last data chunk, it is recommended to check if the `delta.tool_calls` field exists to determine if the current response includes a tool call; * During streaming output, `delta.content` will be output first, followed by `delta.tool_calls`, so you must wait until `delta.content` has finished outputting before you can determine and identify `tool_calls`; * During streaming output, we will specify the `tool_call.id` and `tool_call.function.name` in the initial data chunk, and only `tool_call.function.arguments` will be output in subsequent chunks; * During streaming output, if Kimi returns multiple `tool_calls` at once, we will use an additional field called `index` to indicate the index of the current `tool_call`, so that you can correctly concatenate the `tool_call.function.arguments` parameters. We use a code example from the streaming output section (without using the SDK) to illustrate how to do this: ```python theme={null} import os import json import httpx # We use the httpx library to make our HTTP requests tools = [ { "type": "function", # The type field is fixed as "function" "function": { # When type is "function", use the function field to define the specific function content "name": "search", # The name of the function, please use English letters, numbers, hyphens, and underscores "description": """ Search the internet for content using a search engine. When your knowledge cannot answer the user's question or the user requests an online search, call this tool. Extract the content the user wants to search from the conversation as the value of the query parameter. The search results include the title of the website, the website's address (URL), and a brief introduction to the website. """, # Description of the function, explaining its specific role and usage scenarios to help the Kimi large language model choose the right functions "parameters": { # Use the parameters field to define the parameters the function accepts "type": "object", # Always use type: object to make the Kimi large language model generate a JSON Object parameter "required": ["query"], # Use the required field to tell the Kimi large language model which parameters are mandatory "properties": { # Specific parameter definitions in properties, you can define multiple parameters "query": { # Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", # Use type to define the parameter type "description": """ The content the user wants to search for, extracted from the user's question or chat context. """ # Use description to help the Kimi large language model generate parameters more effectively } } } } }, ] header = { "Content-Type": "application/json", "Authorization": f"Bearer {os.environ.get('MOONSHOT_API_KEY')}", } data = { "model": "kimi-k3", "messages": [ {"role": "user", "content": "Please search for Context Caching technology online."} ], "stream": True, "tools": tools, # <-- Add tool invocation } # Use httpx to send a chat request to the Kimi large language model and get the response r r = httpx.post("https://api.moonshot.ai/v1/chat/completions", headers=header, json=data) if r.status_code != 200: raise Exception(r.text) data: str # Here, we pre-build a List to store different response messages. Since we set n=2, we initialize the List with 2 elements messages = [{}, {}] # Here, we use the iter_lines method to read the response body line by line for line in r.iter_lines(): # Remove leading and trailing spaces from each line to better handle data blocks line = line.strip() # Next, we need to handle three different cases: # 1. If the current line is empty, it indicates that the previous data block has been received (as mentioned earlier, data blocks are ended with two newline characters). We can deserialize the data block and print the corresponding content; # 2. If the current line is not empty and starts with data:, it indicates the start of a data block transmission. After removing the data: prefix, first check if it is the end marker [DONE]. If not, save the data content to the data variable; # 3. If the current line is not empty but does not start with data:, it means the current line still belongs to the previous data block being transmitted. Append the content of the current line to the end of the data variable; if len(line) == 0: chunk = json.loads(data) # Loop through all choices in each data block to get the message object corresponding to the index for choice in chunk["choices"]: index = choice["index"] message = messages[index] usage = choice.get("usage") if usage: message["usage"] = usage delta = choice["delta"] role = delta.get("role") if role: message["role"] = role content = delta.get("content") if content: if "content" not in message: message["content"] = content else: message["content"] = message["content"] + content # From here, we start processing tool_calls tool_calls = delta.get("tool_calls") # <-- First, check if the data block contains tool_calls if tool_calls: if "tool_calls" not in message: message["tool_calls"] = [] # <-- If it contains tool_calls, initialize a list to store these tool_calls. Note that the list is empty at this point, with a length of 0 for tool_call in tool_calls: tool_call_index = tool_call["index"] # <-- Get the index of the current tool_call if len(message["tool_calls"]) < ( tool_call_index + 1): # <-- Expand the tool_calls list according to the index to access the corresponding tool_call via index message["tool_calls"].extend([{}] * (tool_call_index + 1 - len(message["tool_calls"]))) tool_call_object = message["tool_calls"][tool_call_index] # <-- Access the corresponding tool_call via index tool_call_object["index"] = tool_call_index # The following steps fill in the id, type, and function fields of each tool_call based on the information in the data block # In the function field, there are name and arguments fields. The arguments field will be supplemented by each data block # in the same way as the delta.content field. tool_call_id = tool_call.get("id") if tool_call_id: tool_call_object["id"] = tool_call_id tool_call_type = tool_call.get("type") if tool_call_type: tool_call_object["type"] = tool_call_type tool_call_function = tool_call.get("function") if tool_call_function: if "function" not in tool_call_object: tool_call_object["function"] = {} tool_call_function_name = tool_call_function.get("name") if tool_call_function_name: tool_call_object["function"]["name"] = tool_call_function_name tool_call_function_arguments = tool_call_function.get("arguments") if tool_call_function_arguments: if "arguments" not in tool_call_object["function"]: tool_call_object["function"]["arguments"] = tool_call_function_arguments else: tool_call_object["function"]["arguments"] = tool_call_object["function"][ "arguments"] + tool_call_function_arguments # <-- Supplement the value of the function.arguments field sequentially message["tool_calls"][tool_call_index] = tool_call_object data = "" # Reset data elif line.startswith("data: "): data = line[len("data: "):] # When the data block content is [DONE], it indicates that all data blocks have been sent and the network connection can be disconnected if data == "[DONE]": break else: data = data + "\n" + line # When appending content, add a newline character because this might be intentional line breaks in the data block # After assembling all messages, print their contents separately for index, message in enumerate(messages): print("index:", index) print("message:", json.dumps(message, ensure_ascii=False)) print("") ``` ```js theme={null} const os = require('os'); const axios = require('axios');// Use the axios library to perform HTTP requests const tools = [ { "type": "function", "function": { "name": "search", "description": "Search the internet for content using a search engine.\n\nWhen your knowledge cannot answer the user's question or the user requests an online search, call this tool. Extract the content the user wants to search from the conversation as the value of the query parameter.\nThe search results include the title of the website, the website's address (URL), and a brief introduction to the website.", "parameters": { "type": "object", "required": ["query"], "properties": { "query": { "type": "string", "description": "The content the user wants to search for, extracted from the user's question or chat context." } } } } }, ]; const header = { "Content-Type": "application/json", "Authorization": `Bearer ${process.env.MOONSHOT_API_KEY}` }; const data = { "model": "kimi-k3", "messages": [ {"role": "user", "content": "Please search for Context Caching technology online."} ], "stream": true, "tools": tools, "tool_choice": "auto" }; axios.post("https://api.moonshot.ai/v1/chat/completions", data, { headers: header, responseType: 'stream' }).then(response => { if (response.status !== 200) { throw new Error(response.text); } let data = ""; let messages = [{}, {}]; response.data.on('data', chunk => { let line = chunk.toString().trim(); if (line === "") { let chunk = JSON.parse(data); for (let choice of chunk.choices) { let index = choice.index; let message = messages[index]; let usage = choice.usage; if (usage) message.usage = usage; let delta = choice.delta; let role = delta.role; if (role) message.role = role; let content = delta.content; if (content) message.content = (message.content || "") + content; let tool_calls = delta.tool_calls; if (tool_calls) { if (!message.tool_calls) message.tool_calls = []; for (let tool_call of tool_calls) { let tool_call_index = tool_call.index; while (message.tool_calls.length < tool_call_index + 1) { message.tool_calls.push({}); } let tool_call_object = message.tool_calls[tool_call_index]; tool_call_object.index = tool_call_index; let tool_call_id = tool_call.id; if (tool_call_id) tool_call_object.id = tool_call_id; let tool_call_type = tool_call.type; if (tool_call_type) tool_call_object.type = tool_call_type; let tool_call_function = tool_call.function; if (tool_call_function) { if (!tool_call_object.function) tool_call_object.function = {}; let tool_call_function_name = tool_call_function.name; if (tool_call_function_name) tool_call_object.function.name = tool_call_function_name; let tool_call_function_arguments = tool_call_function.arguments; if (tool_call_function_arguments) { if (!tool_call_object.function.arguments) { tool_call_object.function.arguments = tool_call_function_arguments; } else { tool_call_object.function.arguments = tool_call_object.function.arguments + tool_call_function_arguments; } } } message.tool_calls[tool_call_index] = tool_call_object; } } } data = ""; // Reset data } else if (line.startsWith("data: ")) { data = line.substring(6); } else { data = data + "\n" + line; } }); response.data.on('end', () => { for (let index = 0; index < messages.length; index++) { console.log("index:", index); console.log("message:", JSON.stringify(messages[index], null, 4)); console.log(""); } }); }).catch(error => { console.error("Request failed:", error); }); ``` Below is an example of handling `tool_calls` in streaming output using the openai SDK: ```python theme={null} import os import json from openai import OpenAI client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1", ) tools = [ { "type": "function", # The agreed-upon field type, currently supports function as a value "function": { # When type is function, use the function field to define the specific function content "name": "search", # The name of the function, please use English letters, numbers, plus hyphens and underscores as the function name "description": """ Search for content on the internet using a search engine. When your knowledge cannot answer the user's question, or the user requests you to perform an online search, call this tool. Please extract the content the user wants to search from the conversation with the user as the value of the query parameter. The search results include the title of the website, the website's address (URL), and the website's description. """, # The introduction of the function, write the specific function here and its usage scenarios so that the Kimi large language model can correctly choose which functions to use "parameters": { # Use the parameters field to define the parameters accepted by the function "type": "object", # Fixed use type: object to make the Kimi large language model generate a JSON Object parameter "required": ["query"], # Use the required field to tell the Kimi large language model which parameters are required "properties": { # The properties are the specific parameter definitions, you can define multiple parameters "query": { # Here, the key is the parameter name, and the value is the specific definition of the parameter "type": "string", # Use type to define the parameter type "description": """ The content the user is searching for, please extract it from the user's question or chat context. """ # Use description to describe the parameter so that the Kimi large language model can better generate the parameter } } } } }, ] completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "user", "content": "Please search for Context Caching technology online."} ], stream=True, tools=tools, # <-- Add tool invocation ) # Here, we pre-build a List to store different response messages, since we set n=2, we initialize the List with 2 elements messages = [{}, {}] for chunk in completion: # Loop through all the choices in each data chunk and get the message object corresponding to the index for choice in chunk.choices: index = choice.index message = messages[index] delta = choice.delta role = delta.role if role: message["role"] = role content = delta.content if content: if "content" not in message: message["content"] = content else: message["content"] = message["content"] + content # From here, we start processing tool_calls tool_calls = delta.tool_calls # <-- First check if the data chunk contains tool_calls if tool_calls: if "tool_calls" not in message: message["tool_calls"] = [] # <-- If it contains tool_calls, we initialize a list to save these tool_calls, note that the list is empty at this time with a length of 0 for tool_call in tool_calls: tool_call_index = tool_call.index # <-- Get the index of the current tool_call if len(message["tool_calls"]) < ( tool_call_index + 1): # <-- Expand the tool_calls list according to the index so that we can access the corresponding tool_call via the subscript message["tool_calls"].extend([{}] * (tool_call_index + 1 - len(message["tool_calls"]))) tool_call_object = message["tool_calls"][tool_call_index] # <-- Access the corresponding tool_call via the subscript tool_call_object["index"] = tool_call_index # The following steps are to fill in the id, type, and function fields of each tool_call based on the information in the data chunk # In the function field, there are name and arguments fields, the arguments field will be supplemented by each data chunk # Sequentially, just like the delta.content field. tool_call_id = tool_call.id if tool_call_id: tool_call_object["id"] = tool_call_id tool_call_type = tool_call.type if tool_call_type: tool_call_object["type"] = tool_call_type tool_call_function = tool_call.function if tool_call_function: if "function" not in tool_call_object: tool_call_object["function"] = {} tool_call_function_name = tool_call_function.name if tool_call_function_name: tool_call_object["function"]["name"] = tool_call_function_name tool_call_function_arguments = tool_call_function.arguments if tool_call_function_arguments: if "arguments" not in tool_call_object["function"]: tool_call_object["function"]["arguments"] = tool_call_function_arguments else: tool_call_object["function"]["arguments"] = tool_call_object["function"][ "arguments"] + tool_call_function_arguments # <-- Sequentially supplement the value of the function.arguments field message["tool_calls"][tool_call_index] = tool_call_object # After assembling all messages, we print their contents separately for index, message in enumerate(messages): print("index:", index) print("message:", json.dumps(message, ensure_ascii=False)) print("") ``` ```js theme={null} const os = require('os'); const openai = require('openai'); // Need to install the openai library const client = new openai.OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1" }); const tools = [ { "type": "function", "function": { "name": "search", "description": "Search for content on the internet using a search engine.\n\nCall this tool when your knowledge cannot answer the user's question, or when the user requests an online search. Extract the content the user wants to search for from the conversation and use it as the value of the query parameter.\nThe search results include the website title, address (URL), and a brief description of the website.", "parameters": { "type": "object", "required": ["query"], "properties": { "query": { "type": "string", "description": "The content the user wants to search for, extracted from the user's question or chat context." } } } } }, ]; async function main() { const response = await client.chat.completions.create({ model: "kimi-k3", messages: [ { "role": "user", "content": "Please search for Context Caching technology online." } ], stream: true, tools: tools, tool_choice: "auto" }); let messages = [{}, {}]; let data = ''; for await (const chunk of response) { for (const choice of chunk.choices) { const index = choice.index; const message = messages[index]; const delta = choice.delta; const role = delta.role; if (role) message.role = role; const content = delta.content; if (content) message.content = (message.content || "") + content; const tool_calls = delta.tool_calls; if (tool_calls) { if (!message.tool_calls) message.tool_calls = []; for (const tool_call of tool_calls) { const tool_call_index = tool_call.index; if (message.tool_calls.length < tool_call_index + 1) { for (let i = message.tool_calls.length; i < tool_call_index + 1; i++) { message.tool_calls.push({}); } } const tool_call_object = message.tool_calls[tool_call_index]; tool_call_object.index = tool_call_index; const tool_call_id = tool_call.id; if (tool_call_id) tool_call_object.id = tool_call_id; const tool_call_type = tool_call.type; if (tool_call_type) tool_call_object.type = tool_call_type; const tool_call_function = tool_call.function; if (tool_call_function) { if (!tool_call_object.function) tool_call_object.function = {}; const tool_call_function_name = tool_call_function.name; if (tool_call_function_name) tool_call_object.function.name = tool_call_function_name; const tool_call_function_arguments = tool_call_function.arguments; if (tool_call_function_arguments) { if (!tool_call_object.function.arguments) { tool_call_object.function.arguments = tool_call_function_arguments; } else { tool_call_object.function.arguments += tool_call_function_arguments; } } } message.tool_calls[tool_call_index] = tool_call_object; } } } } for (let index = 0; index < messages.length; index++) { console.log("index:", index); console.log("message:", JSON.stringify(messages[index], null, 2)); console.log(""); } } main().catch(console.error); ``` ## Use tool\_calls Instead of function\_call `tool_calls` evolved from function calls (`function_call`), and `function_call` is a subset of `tool_calls` — in certain contexts, or when reading compatibility code, you can treat the two as equivalent. Since OpenAI has marked `function_call` and related parameters (such as `functions`) as "deprecated", our API will no longer support `function_call`; use `tool_calls` instead. Compared to `function_call`, `tool_calls` has the following advantages: * It supports parallel calls. The Kimi large language model can return multiple `tool_calls` at once. You can use concurrency in your code to call these `tool_call` simultaneously, reducing time consumption; * For `tool_calls` that have no dependencies, the Kimi large language model will also tend to call them in parallel. Compared to the original sequential calls of `function_call`, this reduces token consumption to some extent; ## Notes * When `finish_reason=tool_calls`, `message.content` is occasionally not empty: it is usually the Kimi large language model explaining which tools it needs to call and why. If your tool call process takes a long time, or a single turn requires several sequential tool calls, this descriptive message can reduce the anxiety or dissatisfaction users feel while waiting, and helps them understand the tool call flow and intervene and correct in time (for example, terminating an incorrect tool call, or correcting the model's tool selection with a prompt in the next turn); * The content in the `tools` parameter is also counted in the total Tokens. Please ensure that the total number of Tokens in `tools` and `messages` does not exceed the model's context window size. ### Keep Every tool\_call Matched to a tool Message In tool call scenarios, messages are no longer a simple alternation of `system` / `user` / `assistant`: ```text theme={null} system: ... user: ... assistant: ... user: ... assistant: ... ``` Instead, they look like this: ```text theme={null} system: ... user: ... assistant: ... tool: ... tool: ... assistant: ... ``` When the Kimi large language model generates `tool_calls`, make sure every `tool_call` has a corresponding message with `role=tool`, and that this message carries the correct `tool_call_id`: if the number of `role=tool` messages does not match the number of `tool_calls`, an error will occur; likewise, if the `tool_call_id` in a `role=tool` message cannot be matched with the `tool_call.id` in `tool_calls`, an error will occur. ### Troubleshoot the tool\_call\_id not found Error If you encounter the `tool_call_id not found` error, it may be because you did not add the `role=assistant` message returned by the Kimi API to the messages list. The correct message sequence should look like this: ```text theme={null} system: ... user: ... assistant: ... # <-- Perhaps you did not add this assistant message to the messages list tool: ... tool: ... assistant: ... ``` You can avoid the `tool_call_id not found` error by executing `messages.append(message)` each time you receive a return value from the Kimi API, to add the message returned by the Kimi API to the messages list. *Note: Assistant messages added to the messages list before the `role=tool` message must fully include the `tool_calls` field and its values returned by the Kimi API. We recommend directly adding the `choice.message` returned by the Kimi API to the messages list "as is" to avoid potential errors.* # Use Kimi K3 in Hermes Agent Source: https://platform.kimi.ai/docs/guide/use-kimi-in-hermes-agent Install Hermes Agent, connect it to the Global Kimi Open Platform, and enable Kimi K3 text, image, and video understanding. [Hermes Agent](https://github.com/nousresearch/hermes-agent) is an open-source AI agent from Nous Research. It provides persistent memory, tool use, and interfaces including the CLI, Telegram, Discord, Slack, and WhatsApp. This guide connects Hermes Agent to the Global Kimi Open Platform and configures `kimi-k3` with its 1M-token context window and native visual understanding. The API Base URL in Hermes is `https://api.moonshot.ai/v1`; the actual Chat Completions endpoint is `https://api.moonshot.ai/v1/chat/completions`. The screenshots in this guide use Hermes Agent v0.18.2. Menu labels may change in later releases, but continue to use a Global Kimi Open Platform API key, `kimi-k3`, and `https://api.moonshot.ai/v1`. ## Prerequisites Complete these prerequisites before you start. Follow the linked official guides for installation and account setup. Install or update Hermes Agent with the official instructions. Create a Global Kimi Open Platform API key and keep it private. Confirm that the account has available balance, then check limits, budgets, and organization settings. You need available balance to call the API. See [Rate Limits](/docs/pricing/limits) for the current tiers. If your organization uses an IP whitelist, follow [Organization Best Practices](/docs/guide/org-best-practice) and add the public egress IPv4 address of the computer or gateway that calls the API. ## Step 1: Select the Global Kimi provider and K3 Run the model configuration wizard: ```bash theme={null} hermes model ``` Select **Kimi / Moonshot** in the provider list. Select Kimi and Moonshot In Hermes Agent v0.18.2, select the **first submenu item**, currently labeled **Kimi / Kimi Coding Plan**. This is also the built-in entry for the Global Moonshot API. When prompted for `KIMI_API_KEY`, paste the key created at `platform.kimi.ai`; Hermes masks it while you type. Select the first Kimi submenu item For a Global Open Platform key, Hermes should print: ```text theme={null} Using Moonshot endpoint → https://api.moonshot.ai/v1 ``` Confirm the Global Moonshot endpoint If `kimi-k3` is not in the model list, select **Enter custom model name**, then enter `kimi-k3`. Do not add a `moonshot/` prefix. Choose to enter a custom model name Do not put the API key in `config.yaml`, command arguments, screenshots, or a Git repository. This guide uses a Global Kimi Open Platform key; do not substitute a key from another Kimi service or region. ## Step 2: Add the complete K3 configuration Run: ```bash theme={null} hermes config edit ``` Append the `kimi-k3-global` item below to the existing `custom_providers` list, then merge the remaining fields into their corresponding top-level sections. If a top-level section already exists, update its fields instead of creating a duplicate. Preserve unrelated settings and existing custom providers. ```yaml theme={null} custom_providers: - name: kimi-k3-global base_url: https://api.moonshot.ai/v1 key_env: KIMI_API_KEY api_mode: chat_completions model: kimi-k3 extra_body: reasoning_effort: max models: kimi-k3: context_length: 1048576 supports_vision: true model: provider: custom:kimi-k3-global default: kimi-k3 context_length: 1048576 supports_vision: true agent: reasoning_effort: max auxiliary: vision: provider: main model: kimi-k3 extra_body: reasoning_effort: max ``` This configuration explicitly enables Kimi K3's currently supported maximum reasoning effort, 1M-token context window, and native image understanding. It also routes video analysis through the same K3 main provider. `reasoning_effort` is a top-level Chat Completions request field that supports `low`, `high`, and `max` with `max` as the default; the examples use `max`. See the [Model Parameter Reference](/docs/api/models-overview). Set `base_url` to `https://api.moonshot.ai/v1` only. Hermes uses an OpenAI-compatible client and appends `/chat/completions` automatically. Do not put the complete request endpoint in `base_url`. ## Step 3: Enable and use the official video tool Run: ```bash theme={null} hermes tools enable video ``` Hermes can now call its official `video_analyze` tool. For a local video, the tool reads the complete file, encodes it as `data:video/...;base64,...`, and sends it to `kimi-k3` in one `video_url` content block. This path does not extract frames locally with FFmpeg first. Hermes Agent v0.18.2 supports common formats including MP4, WebM, MOV, AVI, MKV, and MPEG, with an approximately 50 MB Base64 video payload limit. This limit comes from the Hermes client's video tool hard cap, not from the Kimi API. Trim or compress larger videos before use. Run the enable command only once. For later video analysis, no additional setup or `/video` command is required. In a Hermes chat, provide an absolute video path and explicitly ask it to call `video_analyze`: ```text theme={null} Call video_analyze on /absolute/path/to/demo.mp4, summarize the video, and list three details that can be confirmed from the visuals. ``` You can also provide a directly accessible HTTP or HTTPS video URL. ## Step 4: Start Hermes Agent and use K3 Start a new session so the provider, context, and tool settings are loaded: ```bash theme={null} hermes ``` The status bar should show `kimi-k3` and a `1M` context window. Use Kimi K3 in Hermes Agent ### Use an image Attach a local image with an absolute path: ```text theme={null} /image /absolute/path/to/example.png ``` You can also copy an image to the clipboard, run `/paste`, and then enter your question. Hermes sends the image to `kimi-k3` as native visual input. ## Troubleshooting ### `kimi-k3` is missing from the model list Run `hermes model` again, select **Kimi / Moonshot > Kimi / Kimi Coding Plan**, then choose **Enter custom model name** and enter `kimi-k3`. ### Hermes resolves the wrong endpoint For this guide, Hermes must print `Using Moonshot endpoint → https://api.moonshot.ai/v1`. Use an API key created at `https://platform.kimi.ai/console/api-keys`, then run `hermes model` again. ### The API key is rejected Re-enter the Global Kimi Open Platform key with `hermes model`. If your organization uses an IP whitelist, confirm that the caller's public egress IPv4 address is allowed. ### You receive a 429 error Reduce concurrent requests, wait before retrying, and check the account balance and current [rate-limit tier](/docs/pricing/limits). ### The video tool is not called Confirm that you ran `hermes tools enable video`, explicitly ask Hermes to call `video_analyze`, and use an absolute local path. The encoded payload must remain below approximately 50 MB. For more Kimi API guidance, see [Troubleshooting](/docs/guide/troubleshooting). # Use Kimi in OpenClaw Source: https://platform.kimi.ai/docs/guide/use-kimi-in-openclaw Build a cross-platform AI agent with OpenClaw and the Kimi API. Install OpenClaw and configure your Kimi API key. OpenClaw (formerly Clawdbot and Moltbot) is an open-source, self-hosted AI agent platform that lets you run AI assistants locally. It integrates with WhatsApp, Telegram, Discord, Slack, and Signal to connect large language models to real-world workflows. The platform supports multiple LLM providers, extensible skills, and gives you full control over your data and API keys. The steps below use OpenClaw `2026.7.1` with the official Moonshot Provider and Kimi K3 for Chat Completion. The old K2.5 preset is not the target configuration in this guide. `kimi-k2.5` is no longer available to newly registered users. New users should use the official Moonshot Provider and set `moonshot/kimi-k3` as the default model. Enter API keys only in the local wizard or terminal; never put them in docs, screenshots, repositories, or chat messages. ## Prerequisites Complete these prerequisites before you start. Follow the official installation, source, and account links below; this guide focuses only on configuring Kimi in OpenClaw. Install or update OpenClaw using the official entry point. Review the official repository, releases, and change history. Create and keep private an API key from the international Kimi Open Platform. Your account needs available balance to use Kimi K3. See [Rate Limits](/docs/pricing/limits) for current tiers. If your organization uses an IP whitelist, follow [Organization Best Practices](/docs/guide/org-best-practice) and add the public egress IPv4 address of the calling computer or gateway. ## Set up Kimi K3 Make sure OpenClaw is installed from the prerequisites above, then install or update the official Moonshot Provider: ```bash theme={null} openclaw plugins install @openclaw/moonshot-provider openclaw gateway restart ``` Run onboarding: ```bash theme={null} openclaw onboard --auth-choice moonshot-api-key ``` In the wizard, choose: * **Step 1: Model.auth provider > Choose Moonshot** * **Step 2: Model AI auth method > Choose Kimi API key (.ai)** * **Step 3: Enter Moonshot API Key (.ai) > Enter your international API key** * **Step 4: Default model > Set it to `moonshot/kimi-k3` after onboarding** If K3 is not shown in the wizard, finish authentication first, then run: ```bash theme={null} openclaw models list --provider moonshot openclaw models set moonshot/kimi-k3 ``` If the stable Provider catalog still does not include K3, upgrade the plugin and restart the Gateway: ```bash theme={null} openclaw plugins update @openclaw/moonshot-provider openclaw gateway restart openclaw models list --provider moonshot ``` Screenshots from the old onboarding flow may still show the historical `Moonshot AI (Kimi K2.5)` or `moonshot/kimi-k2.5` preset. This guide no longer uses that preset; follow the text steps and the K3 model selector screenshot below, and do not keep the old default model. If K3 is still missing after the upgrade, add a K3 entry to `models.providers.moonshot.models` in `~/.openclaw/openclaw.json`, keeping existing models: ```json theme={null} { "models": { "mode": "merge", "providers": { "moonshot": { "baseUrl": "https://api.moonshot.ai/v1", "api": "openai-completions", "models": [{ "id": "kimi-k3", "name": "Kimi K3", "reasoning": true, "input": ["text", "image", "video"], "contextWindow": 1048576, "maxTokens": 8192, "thinkingLevelMap": { "off": null, "minimal": "max", "low": "max", "medium": "max", "high": "max", "xhigh": "max", "max": "max" }, "compat": { "maxTokensField": "max_tokens", "supportsUsageInStreaming": false, "requiresStringContent": true, "supportsReasoningEffort": true, "supportedReasoningEfforts": ["minimal", "low", "medium", "high", "xhigh", "max"] } }] } } } } ``` To route image and video inputs through K3 as well, add this to the same configuration: ```json theme={null} { "agents": { "defaults": { "imageModel": "moonshot/kimi-k3" } }, "tools": { "media": { "image": { "models": [{ "type": "provider", "provider": "moonshot", "model": "kimi-k3", "capabilities": ["image"] }] }, "video": { "models": [{ "type": "provider", "provider": "moonshot", "model": "kimi-k3", "capabilities": ["video"] }] } } } } ``` K3 always uses the server-side `max` reasoning setting. `contextWindow` remains 1M; `maxTokens` is set to 8192 as the OpenClaw per-reply limit, so the 1M input window is not mistaken for a 1M output limit. Keep `maxTokensField: "max_tokens"`, `supportsUsageInStreaming: false`, and `requiresStringContent: true` for K3 compatibility: K3 treats `max_completion_tokens`, streaming usage, and array-form pure-text content differently. ## Step 4: Start using OpenClaw After installation, open the Control UI address shown by the onboarding wizard or Gateway output. The bottom of the Control UI should show `kimi-k3 · moonshot`. You can now send messages in the chat page. The model selector screenshot is shown above: K3 chat page ## Troubleshooting ### 401 / Invalid Authentication * Confirm that you are using a Kimi Open Platform API key, not a Kimi Code key. * Use `moonshot-api-key` for an international key; use the Chinese edition for a China-region key. * If an old key is already present in the environment, rerun onboarding and enter the international key again. ### The Moonshot authentication option is unavailable Make sure the official Provider is installed and the Gateway has been restarted: ```bash theme={null} openclaw plugins install @openclaw/moonshot-provider openclaw gateway restart ``` ### K3 is missing from the model list Upgrade OpenClaw and the Moonshot Provider. If it is still missing, add the K3 entry from Step 3 and run `openclaw models set moonshot/kimi-k3` again. For more Kimi API guidance, see [Troubleshooting](/docs/guide/troubleshooting). # Build an Agent with Kimi K3 Source: https://platform.kimi.ai/docs/guide/use-kimi-k3-to-setup-agent Build a runnable industry-research agent by combining Kimi K3, the official web-search tool, and a custom tool. Kimi K3 provides reasoning, coding, and tool-calling capabilities for complex tasks. This guide uses an industry-research agent to show how to combine an official web-search tool with one custom tool in a runnable, bounded agent. ## Break Down the Task Split industry research into three stages before choosing tools and writing prompts: 1. **Retrieve**: define the research scope and find current data, company information, and news; 2. **Analyze**: compare sources, identify conflicts, and separate facts, estimates, and inferences; 3. **Deliver**: produce a structured report with a summary, key findings, risks, and sources. This split lets the model plan and make judgments while tools handle retrieval or deterministic operations. Avoid repeating tool behavior in a long system prompt. ## Design Tools This example combines two kinds of tools: * Official `web-search`: retrieves current industry sources. The platform also provides `fetch`, `code-runner`, `excel`, and other tools. See [Official Tools](/docs/guide/use-official-tools) for the full list and Formula API flow; * Custom `build_research_plan`: generates a deterministic scope from a topic, region, and year range, demonstrating how to declare and execute a local function tool. Custom tools describe arguments with JSON Schema. Setting `additionalProperties` to `false` and listing mandatory fields in `required` reduces invalid arguments. When your inventory grows to dozens or hundreds of tools, do not send every schema in every request. Follow [Kimi K3 Tool Calling Best Practices](/docs/guide/kimi-k3-tool-calling-best-practice) and [Dynamically Loaded Tools](/docs/guide/use-dynamic-tool-loading) to retrieve candidates first and load them on demand. ## Design the Prompt Keep the system prompt focused on the role, workflow, and quality boundaries. Leave argument details in the tool schema: ```python theme={null} SYSTEM_PROMPT = """You are an industry research assistant. Define the scope first, retrieve and cross-check information, then write a concise report. Requirements: - Separate confirmed facts, estimates, and inferences; support key findings with multiple sources where possible. - Never fabricate data or sources; state the search scope and gaps when evidence is insufficient. - Include an executive summary, key findings, risks and limitations, and a source list. - Answer in the same language as the user. """ ``` Add constraints incrementally when the business format, compliance requirements, or audience changes. See [Prompt Best Practices](/docs/guide/prompt-best-practice) for more guidance. ## Configure the K3 API Use Python 3.9 or later, and install the OpenAI Python SDK and the HTTP client used for official Formula tools: ```bash theme={null} python3 -m pip install --upgrade openai httpx export MOONSHOT_API_KEY="YOUR_API_KEY" ``` The Global endpoint is `https://api.moonshot.ai/v1`, and the model is `kimi-k3`. The API key is read only from the `MOONSHOT_API_KEY` environment variable. Kimi K3 always reasons, and its reasoning effort is configured with the top-level `reasoning_effort` request field, which supports `"low"` / `"high"` / `"max"` (default `"max"`). A tool loop must append the complete assistant message returned by the SDK to `messages`. Copying only `content` and `tool_calls` drops any returned `reasoning_content` and breaks the context needed by later tool calls. For parameters and model-specific behavior, see [Thinking Mode](/docs/guide/use-thinking-models), [Reasoning Effort](/docs/guide/use-reasoning-effort), and the [Model Parameter Reference](/docs/api/models-overview). ## Complete Agent Loop Save the following code as `agent.py`. It loads the `web-search` declaration dynamically, executes the custom tool locally, and handles tool calls in a loop capped at eight rounds. ```python theme={null} import asyncio import json import os import httpx from openai import AsyncOpenAI BASE_URL = "https://api.moonshot.ai/v1" MODEL = "kimi-k3" MAX_TOOL_ROUNDS = 8 SYSTEM_PROMPT = """You are an industry research assistant. Define the scope first, retrieve and cross-check information, then write a concise report. Requirements: - Separate confirmed facts, estimates, and inferences; support key findings with multiple sources where possible. - Never fabricate data or sources; state the search scope and gaps when evidence is insufficient. - Include an executive summary, key findings, risks and limitations, and a source list. - Answer in the same language as the user. """ RESEARCH_PLAN_TOOL = { "type": "function", "function": { "name": "build_research_plan", "description": "Build a research plan from an industry topic, region, and time range", "parameters": { "type": "object", "properties": { "topic": { "type": "string", "description": "Industry or subject to research", }, "region": { "type": "string", "enum": ["China", "Global", "United States", "Europe"], "description": "Region to research", }, "start_year": { "type": "integer", "description": "First year in the research period", }, "end_year": { "type": "integer", "description": "Last year in the research period", }, }, "required": ["topic", "region", "start_year", "end_year"], "additionalProperties": False, }, }, } def build_research_plan( topic: str, region: str, start_year: int, end_year: int ) -> str: """Build a deterministic scope as a minimal custom-tool example.""" if start_year > end_year: return json.dumps({"error": "start_year must not exceed end_year"}) plan = { "topic": topic, "region": region, "period": f"{start_year}-{end_year}", "dimensions": [ "market size and growth", "value chain and major companies", "technology trends", "policy and risks", ], "search_queries": [ f"{region} {topic} market size {start_year} {end_year}", f"{region} {topic} major companies technology trends", f"{region} {topic} policy risks", ], } return json.dumps(plan) class IndustryResearchAgent: def __init__(self) -> None: api_key = os.environ["MOONSHOT_API_KEY"] self.openai = AsyncOpenAI(api_key=api_key, base_url=BASE_URL) self.http = httpx.AsyncClient( base_url=BASE_URL, headers={"Authorization": f"Bearer {api_key}"}, timeout=60.0, ) async def load_formula( self, formula_uri: str ) -> tuple[list[dict], dict[str, str]]: response = await self.http.get(f"/formulas/{formula_uri}/tools") response.raise_for_status() tools = response.json().get("tools", []) if not tools: raise RuntimeError(f"Formula {formula_uri} returned no tools") tool_to_formula = { tool["function"]["name"]: formula_uri for tool in tools if tool.get("type") == "function" and tool.get("function") } if not tool_to_formula: raise RuntimeError(f"Formula {formula_uri} returned no callable function tools") return tools, tool_to_formula async def call_formula( self, formula_uri: str, name: str, arguments: dict ) -> str: response = await self.http.post( f"/formulas/{formula_uri}/fibers", json={"name": name, "arguments": json.dumps(arguments)}, ) response.raise_for_status() fiber = response.json() context = fiber.get("context", {}) if fiber.get("status") == "succeeded": result = context.get("output") or context.get("encrypted_output") or "" else: result = fiber.get("error") or context.get("error") or "Unknown tool error" if isinstance(result, str): return result return json.dumps(result) async def run(self, question: str) -> str: official_tools, tool_to_formula = await self.load_formula( "moonshot/web-search:latest" ) tools = [RESEARCH_PLAN_TOOL, *official_tools] messages = [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": question}, ] for _ in range(MAX_TOOL_ROUNDS): response = await self.openai.chat.completions.create( model=MODEL, messages=messages, tools=tools, max_completion_tokens=8192, ) choice = response.choices[0] message = choice.message # Append the complete SDK message, preserving reasoning_content and tool_calls. messages.append(message) if choice.finish_reason == "length": raise RuntimeError( "The model output was truncated by max_completion_tokens; " "increase the limit or shorten tool results" ) if not message.tool_calls: if choice.finish_reason != "stop": raise RuntimeError( f"Unexpected finish reason: {choice.finish_reason}" ) if not message.content: raise RuntimeError("The model returned no final report") return message.content for tool_call in message.tool_calls: name = tool_call.function.name try: arguments = json.loads(tool_call.function.arguments or "{}") if not isinstance(arguments, dict): raise ValueError("Tool arguments must be a JSON object") if name == "build_research_plan": result = build_research_plan(**arguments) elif name in tool_to_formula: result = await self.call_formula( tool_to_formula[name], name, arguments ) else: raise ValueError(f"Unknown tool: {name}") except Exception as exc: result = json.dumps( {"error": f"{type(exc).__name__}: {exc}"} ) # Return every result and continue other calls even if one tool fails. messages.append( { "role": "tool", "tool_call_id": tool_call.id, "content": result, } ) raise RuntimeError(f"Tool calls exceeded {MAX_TOOL_ROUNDS} rounds") async def close(self) -> None: try: await self.openai.close() finally: await self.http.aclose() async def main() -> None: agent = IndustryResearchAgent() try: report = await agent.run( "Research the development of China's humanoid-robot industry from 2024 to 2026" ) print(report) finally: await agent.close() if __name__ == "__main__": asyncio.run(main()) ``` Two context details in the loop are mandatory: 1. `messages.append(message)` appends the complete assistant message and preserves K3's `reasoning_content`; 2. every tool message uses the matching `tool_call_id`, so the model can associate each result with its call. The example fails immediately on `finish_reason="length"` instead of treating truncated output as a final report. If one tool call fails, its error is returned as the matching tool result while the loop continues processing the other calls in that round. `MAX_TOOL_ROUNDS` prevents endless tool calls and avoids the growing call stack of a recursive implementation. ## Run and Troubleshoot Run the example: ```bash theme={null} python3 agent.py ``` Replace the question in `main()` with your target industry, region, and years. The final response should contain a summary, key findings, risks, and sources instead of a dump of raw tool output. Common issues: * **`finish_reason` is `length`**: the example raises `RuntimeError`; increase `max_completion_tokens` using the [Model Parameter Reference](/docs/api/models-overview), or shorten tool results; * **Maximum tool rounds reached**: check for overlapping tool descriptions and ambiguous results, then narrow the task; * **A tool repeatedly returns argument errors**: check the function schema, `required`, `enum`, and `additionalProperties`; do not replace argument constraints with prompt text; * **A later tool round fails**: confirm that you append the complete SDK assistant message and preserve the correct `tool_call_id` on every result; * **An official tool request fails**: verify the endpoint, API key, Formula URI, and tool availability. See [Official Tools](/docs/guide/use-official-tools) for the full API flow. ## Custom Tools and Optimization The included `build_research_plan` example covers the complete path: declare a schema, dispatch locally, return JSON, and attach the matching `tool_call_id`. To connect a database, internal search service, or file generator, replace the function implementation while keeping its schema and return shape aligned. For further optimization: * expose only tools needed for the current task and avoid overlapping tools; * validate tool inputs in your application and return structured errors so the model can correct arguments; * use [Dynamically Loaded Tools](/docs/guide/use-dynamic-tool-loading) for large inventories; * before using `tool_choice`, reasoning effort, or related controls, check [Kimi K3 Tool Calling Best Practices](/docs/guide/kimi-k3-tool-calling-best-practice) and the [Model Parameter Reference](/docs/api/models-overview). # Configure Kimi Vision Models Source: https://platform.kimi.ai/docs/guide/use-kimi-vision-model Build image and video inputs for Kimi multimodal models using base64, URLs, file uploads, and multi-image conversations. The Kimi Vision Model (including `kimi-k3` / `moonshot-v1-8k-vision-preview` / `moonshot-v1-32k-vision-preview` / `moonshot-v1-128k-vision-preview` / `kimi-k2.5`/ `kimi-k2.6` / `kimi-k2.7-code` / `kimi-k2.7-code-highspeed` and so on) can understand visual content, including text in the image, colors, and the shapes of objects. The `kimi-k3`, `kimi-k2.6`, `kimi-k2.7-code` and `kimi-k2.7-code-highspeed` models can also understand video content. When you need the model to recognize images or videos, build multimodal requests as described on this page. ## Upload images directly with base64 The following example encodes a local image as base64, passes it to Kimi as an `image_url` message part, and asks a question about the image: ```python theme={null} import os import base64 from openai import OpenAI client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1", ) # Replace kimi.png with the path to the image you want Kimi to recognize image_path = "kimi.png" with open(image_path, "rb") as f: image_data = f.read() # We use the built-in base64.b64encode function to encode the image into a base64 formatted image_url image_url = f"data:image/{os.path.splitext(image_path)[1].lstrip('.')};base64,{base64.b64encode(image_data).decode('utf-8')}" completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "system", "content": "You are Kimi."}, { "role": "user", # Note here, the content has changed from the original str type to a list. This list contains multiple parts, with the image (image_url) being one part and the text (text) being another part. "content": [ { "type": "image_url", # <-- Use the image_url type to upload the image, the content is the base64 encoded image "image_url": { "url": image_url, }, }, { "type": "text", "text": "Describe the content of the image.", # <-- Use the text type to provide text instructions, such as "Describe the content of the image" }, ], }, ], ) print(completion.choices[0].message.content) ``` When using a Vision model, `message.content` must be an `array[object]` (that is, a JSON array). **Do not** serialize the JSON array and put it into `message.content` as a `string`. This is a non-standard format and is not guaranteed to be processed as visual input; behavior may vary across models or versions. Always use the array format shown below. The correct format — `content` is a JSON array of multiple parts: ```json theme={null} { "model": "kimi-k3", "messages": [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI, who excels in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You will reject any questions related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated into other languages." }, { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAGAAAABhCAYAAAApxKSdAAAACXBIWXMAACE4AAAhOAFFljFgAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAUUSURBVHgB7Z29bhtHFIWPHQN2J7lKqnhYpYvpIukCbJEAKQJEegLReYFIT0DrCSI9QEDqCSIDaQIEIOukiJwyza5SJWlId3FFz+HuGmuSSw6p+dlZ3g84luhdUeI9M3fmziyXgBCUe/DHYY0Wj/tgWmjV42zFcWe4MIBBPNJ6qqW0uvAbXFvQgKzQK62bQhkaCIPc10q1Zi3XH1o/IG9cwUm0RogrgDY1KmLgHYX9DvyiBvDYI77XmiD+oLlQHw7hIDoCMBOt1U9w0BsU9mOAtaUUFk3oQoIfzAQFCf5dNMEdTFCQ4NtQih1NSIGgf3ibxOJt5UrAB1gNK72vIdjiI61HWr+YnNxDXK0rJiULsV65GJeiIescLSTTeobKSutiCuojX8kU3MBx4I3WeNVBBRl4fWiCyoB8v2JAAkk9PmDwT8sH1TEghRjgC27scCx41wO43KAg+ILxTvhNaUACwTc04Z0B30LwzTzm5Rjw3sgseIG1wGMawMBPIOQcqvzrNIMHOg9Q5KK953O90/rFC+BhJRH8PQZ+fu7SjC7HAIV95yu99vjlxfvBJx8nwHd6IfNJAkccOjHg6OgIs9lsra6vr2GTNE03/k7q8HAhyJ/2gM9O65/4kT7/mwEcoZwYsPQiV3BwcABb9Ho9KKU2njccDjGdLlxx+InBBPBAAR86ydRPaIC9SASi3+8bnXd+fr78nw8NJ39uDJjXAVFPP7dp/VmWLR9g6w6Huo/IOTk5MTpvZesn/93AiP/dXCwd9SyILT9Jko3n1bZ+8s8rGPGvoVHbEXcPMM39V1dX9Qd/19PPNxta959D4HUGF0RrAFs/8/8mxuPxXLUwtfx2WX+cxdivZ3DFA0SKldZPuPTAKrikbOlMOX+9zFu/Q2iAQoSY5H7mfeb/tXCT8MdneU9wNNCuQUXZA0ynnrUznyqOcrspUY4BJunHqPU3gOgMsNr6G0B0BpgUXrG0fhKVAaaF1/HxMWIhKgNMcj9Tz82Nk6rVGdav/tJ5eraJ0Wi01XPq1r/xOS8uLkJc6XYnRTMNXdf62eIvLy+jyftVghnQ7Xahe8FW59fBTRYOzosDNI1hJdz0lBQkBflkMBjMU5iL13pXRb8fYAJrB/a2db0oFHthAOEUliaYFHE+aaUBdZsvvFhApyM0idYZwOCvW4JmIWdSzPmidQaYrAGZ7iX4oFUGnJ2dGdUCTRqMozeANQCLsE6nA10JG/0Mx4KmDMbBCjEWR2yxu8LAM98vXelmCA2ovVLCI8EMYODWbpbvCXtTBzQVMSAwYkBgxIDAtNKAXWdGIRADAiMpKDA0IIMQikx6QGDEgMCIAYGRMSAsMgaEhgbcQgjFa+kBYZnIGBCWWzEgLPNBOJ6Fk/aR8Y5ZCvktKwX/PJZ7xoVjfs+4chYU11tK2sE85qUBLyH4Zh5z6QHhGPOf6r2j+TEbcgdFP2RaHX5TrYQlDflj5RXE5Q1cG/lWnhYpReUGKdUewGnRmhvnCJbgmxey8sHiZ8iwF3AsUBBckKHI/SWLq6HsBc8huML4DiK80D6WnBqLzN68UFCmopheYJOVYgcU5FOVbAVfYUcUZGoaLPglCtITdg2+tZUFBTFh2+ArWEYh/7z0WIIQSiM43lt5AWAmWhLHylN4QmkNEXfAbGqEQKsHSfHLYwiSq8AnaAAKeaW3D8VbijwNW5nh3IN9FPI/jnpaPKZi2/SfFuJu4W3x9RqWL+N5C+7ruKpBAgLkAAAAAElFTkSuQmCC" } }, { "type": "text", "text": "Please describe this image." } ] } ] } ``` The invalid format — the array is serialized into a string: ```json theme={null} { "model": "kimi-k3", "messages": [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate responses. You will refuse to answer any questions involving terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated into other languages." }, { "role": "user", "content": "[{\"type\": \"image_url\", \"image_url\": {\"url\": \"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAGAAAABhCAYAAAApxKSdAAAACXBIWXMAACE4AAAhOAFFljFgAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAUUSURBVHgB7Z29bhtHFIWPHQN2J7lKqnhYpYvpIukCbJEAKQJEegLReYFIT0DrCSI9QEDqCSIDaQIEIOukiJwyza5SJWlId3FFz+HuGmuSSw6p+dlZ3g84luhdUeI9M3fmziyXgBCUe/DHYY0Wj/tgWmjV42zFcWe4MIBBPNJ6qqW0uvAbXFvQgKzQK62bQhkaCIPc10q1Zi3XH1o/IG9cwUm0RogrgDY1KmLgHYX9DvyiBvDYI77XmiD+oLlQHw7hIDoCMBOt1U9w0BsU9mOAtaUUFk3oQoIfzAQFCf5dNMEdTFCQ4NtQih1NSIGgf3ibxOJt5UrAB1gNK72vIdjiI61HWr+YnNxDXK0rJiULsV65GJeiIescLSTTeobKSutiCuojX8kU3MBx4I3WeNVBBRl4fWiCyoB8v2JAAkk9PmDwT8sH1TEghRjgC27scCx41wO43KAg+ILxTvhNaUACwTc04Z0B30LwzTzm5Rjw3sgseIG1wGMawMBPIOQcqvzrNIMHOg9Q5KK953O90/rFC+BhJRH8PQZ+fu7SjC7HAIV95yu99vjlxfvBJx8nwHd6IfNJAkccOjHg6OgIs9lsra6vr2GTNE03/k7q8HAhyJ/2gM9O65/4kT7/mwEcoZwYsPQiV3BwcABb9Ho9KKU2njccDjGdLlxx+InBBPBAAR86ydRPaIC9SASi3+8bnXd+fr78nw8NJ39uDJjXAVFPP7dp/VmWLR9g6w6Huo/IOTk5MTpvZesn/93AiP/dXCwd9SyILT9Jko3n1bZ+8s8rGPGvoVHbEXcPMM39V1dX9Qd/19PPNxta959D4HUGF0RrAFs/8/8mxuPxXLUwtfx2WX+cxdivZ3DFA0SKldZPuPTAKrikbOlMOX+9zFu/Q2iAQoSY5H7mfeb/tXCT8MdneU9wNNCuQUXZA0ynnrUznyqOcrspUY4BJunHqPU3gOgMsNr6G0B0BpgUXrG0fhKVAaaF1/HxMWIhKgNMcj9Tz82Nk6rVGdav/tJ5eraJ0Wi01XPq1r/xOS8uLkJc6XYnRTMNXdf62eIvLy+jyftVghnQ7Xahe8FW59fBTRYOzosDNI1hJdz0lBQkBflkMBjMU5iL13pXRb8fYAJrB/a2db0oFHthAOEUliaYFHE+aaUBdZsvvFhApyM0idYZwOCvW4JmIWdSzPmidQaYrAGZ7iX4oFUGnJ2dGdUCTRqMozeANQCLsE6nA10JG/0Mx4KmDMbBCjEWR2yxu8LAM98vXelmCA2ovVLCI8EMYODWbpbvCXtTBzQVMSAwYkBgxIDAtNKAXWdGIRADAiMpKDA0IIMQikx6QGDEgMCIAYGRMSAsMgaEhgbcQgjFa+kBYZnIGBCWWzEgLPNBOJ6Fk/aR8Y5ZCvktKwX/PJZ7xoVjfs+4chYU11tK2sE85qUBLyH4Zh5z6QHhGPOf6r2j+TEbcgdFP2RaHX5TrYQlDflj5RXE5Q1cG/lWnhYpReUGKdUewGnRmhvnCJbgmxey8sHiZ8iwF3AsUBBckKHI/SWLq6HsBc8huML4DiK80D6WnBqLzN68UFCmopheYJOVYgcU5FOVbAVfYUcUZGoaLPglCtITdg2+tZUFBTFh2+ArWEYh/7z0WIIQSiM43lt5AWAmWhLHylN4QmkNEXfAbGqEQKsHSfHLYwiSq8AnaAAKeaW3D8VbijwNW5nh3IN9FPI/jnpaPKZi2/SfFuJu4W3x9RqWL+N5C+7ruKpBAgLkAAAAAElFTkSuQmCC\"}}, {\"type\": \"text\", \"text\": \"Please describe this image\"}]" } ] } ``` ## Reference uploaded images or videos by file ID Since video files are often larger, you can first upload images or videos to Moonshot and then reference them via file ID — see [Image Understanding Upload](/docs/api/files-upload) for the upload API. The following example uploads a video file and asks the model to describe it through a `video_url` using the `ms://` protocol: ```python theme={null} import os from pathlib import Path from openai import OpenAI client = OpenAI( api_key=os.environ.get("MOONSHOT_API_KEY"), base_url="https://api.moonshot.ai/v1", ) # Here, you need to replace video.mp4 with the path to the image or video you want Kimi to recognize video_path = "video.mp4" file_object = client.files.create(file=Path(video_path), purpose="video") # Upload video to Moonshot completion = client.chat.completions.create( model="kimi-k3", messages=[ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI, who excels in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You will refuse to answer any questions involving terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated into other languages." }, { "role": "user", "content": [ { "type": "video_url", "video_url": { "url": f"ms://{file_object.id}" # Note this is ms:// instead of base64-encoded image } }, { "type": "text", "text": "Please describe this video" } ] } ] ) print(completion.choices[0].message.content) ``` Note that in the above example, the format of `video_url.url` is `ms://`, where ms is short for moonshot storage — Moonshot's internal protocol for referencing files. ## Supported image and video formats ### Images Images support the following formats (MIME types): * image/jpeg * image/png * image/gif * image/webp * image/bmp * image/heic * image/heif Animated GIF and WebP images are also passed in via `image_url`, but the underlying system may decode them as videos, and token consumption is calculated as video accordingly. ### SVG SVG is not supported as image input: uploading an SVG with `purpose="image"`, or passing one through `image_url` (including base64-encoded), will be rejected. To have the model interpret SVG content, include the SVG source (XML text) directly as text in `messages`. ### Videos Videos support the following formats (MIME types): * video/mp4 * video/mpeg * video/mov * video/avi * video/x-flv * video/mpg * video/webm * video/wmv * video/3gpp ## Estimate token usage and costs * Images and videos use dynamic token calculation: you can obtain the token consumption of a request containing images or videos through the [estimate tokens API](/docs/api/estimate) before starting the understanding process; * Generally speaking, the higher the image resolution, the more tokens it consumes. Videos are composed of several key frames — the more key frames and the higher the resolution, the more tokens are consumed; * The Vision model follows the same pricing model as the `moonshot-v1` series, with costs based on the total tokens used for model inference. For token pricing, see [Model Inference Pricing](/docs/pricing/chat-k27-code). ## Keep image and video resolution within limits We recommend that image resolution does not exceed 4k (4096×2160), and video resolution does not exceed FHD (1920×1080). Resolutions higher than recommended will only cost more time processing the input without improving model understanding performance. ## Choose between base64 and file upload * Due to our overall request body size limitations, very large videos must be processed using the file upload method for visual understanding; * For images or videos that need to be referenced multiple times, we recommend using the file upload method for visual understanding; * Regarding file upload limitations, please refer to the [File Upload](/docs/api/files-upload) documentation. ## Feature support and limitations The Vision model supports the following features: * Multi-turn conversations * Streaming output * Tool invocation * JSON Mode * Partial Mode The following features are not supported or only partially supported: * URL-formatted images: Not supported, currently only supports base64-encoded image content and images/videos uploaded via file ID Other limitations: * Image quantity: The Vision model has no limit on the number of images, but ensure that the request body size does not exceed 100M. Different models enforce different constraints on parameters such as `temperature`, `top_p`, and `n`; we recommend using the defaults instead of setting them manually. See [Model Parameter Differences](/docs/api/models-overview) for details. # MoonPalace - Moonshot AI's Kimi API Debugging Tool Source: https://platform.kimi.ai/docs/guide/use-moonpalace Install and use MoonPalace to debug Kimi API conversations, parameters, files, tool calls, and context caching across platforms. MoonPalace (Moon Palace) is an API debugging tool provided by Moonshot AI. It has the following features: * **Cross-platform support**: * [x] Mac * [x] Windows * [x] Linux * **Easy to use**, just replace `base_url` with `http://localhost:9988` after launching to start debugging; * **Captures complete requests**, including the "scene of the accident" when network errors occur; * **Quickly search and view request information** using `request_id` and `chatcmpl_id`; * **One-click export of BadCase structured reporting data**, helping to enhance Kimi's model capabilities; **We recommend using MoonPalace as your API "supplier" during the code writing and debugging phase, so you can quickly identify and locate various issues related to API calls and code writing. For any unexpected outputs from Kimi large language model, you can also export the request details via MoonPalace and submit them to Moonshot AI to improve Kimi large language model.** ## Installation Methods ### Using the `go` Command to Install If you have the `go` toolchain installed, you can run the following command to install MoonPalace: ```shell theme={null} $ go install github.com/MoonshotAI/moonpalace@latest ``` The above command will install the compiled binary file in your `$GOPATH/bin/` directory. Run the `moonpalace` command to check if it has been installed successfully: ```shell theme={null} $ moonpalace MoonPalace is a command-line tool for debugging the Moonshot AI HTTP API. Usage: moonpalace [command] Available Commands: cleanup Cleanup Moonshot AI requests. completion Generate the autocompletion script for the specified shell export export a Moonshot AI request. help Help about any command inspect Inspect the specific content of a Moonshot AI request. list Query Moonshot AI requests based on conditions. start Start the MoonPalace proxy server. Flags: -h, --help help for moonpalace -v, --version version for moonpalace Use "moonpalace [command] --help" for more information about a command. ``` *If you still cannot find the `moonpalace` binary file, try adding the `$GOPATH/bin/` directory to your `$PATH` environment variable.* ### Downloading from the Releases Page You can download the precompiled binary (executable) files from the [Releases](https://github.com/MoonshotAI/moonpalace/releases) page: * moonpalace-linux * moonpalace-macos-amd64 for Intel-based Macs * moonpalace-macos-arm64 for Apple Silicon-based Macs * moonpalace-windows.exe Download the binary (executable) file that matches your platform and place it in a directory that is included in your `$PATH` environment variable. Rename it to `moonpalace` and then grant it executable permissions. ## Usage ### Starting the Service Use the following command to start the MoonPalace proxy server: ```console theme={null} $ moonpalace start --port ``` MoonPalace will start an HTTP server locally, with the `--port` parameter specifying the local port that MoonPalace will listen on. The default value is `9988`. When MoonPalace starts successfully, it will output: ```console theme={null} [MoonPalace] 2024/07/29 17:00:29 MoonPalace Starts {'=>'} change base_url to "http://127.0.0.1:9988/v1" ``` As instructed, replace `base_url` with the displayed address. If you are using the default port, set `base_url=http://127.0.0.1:9988/v1`. If you are using a custom port, replace `base_url` with the displayed address. **Additionally, if you want to always use a debugging `api_key` during debugging, you can use the `--key` parameter when starting MoonPalace to set a default `api_key` for MoonPalace. This way, you don't have to manually set the `api_key` in each request. MoonPalace will automatically add the `api_key` you set with `--key` when requesting the Kimi API.** If you have correctly set `base_url` and successfully called the Kimi API, MoonPalace will output the following information: ```console theme={null} $ moonpalace start --port [MoonPalace] 2024/07/29 17:00:29 MoonPalace Starts {'=>'} change base_url to "http://127.0.0.1:9988/v1" [MoonPalace] 2024/07/29 21:30:53 POST /v1/chat/completions 200 OK [MoonPalace] 2024/07/29 21:30:53 - Request Headers: [MoonPalace] 2024/07/29 21:30:53 - Content-Type: application/json [MoonPalace] 2024/07/29 21:30:53 - Response Headers: [MoonPalace] 2024/07/29 21:30:53 - Content-Type: application/json [MoonPalace] 2024/07/29 21:30:53 - Msh-Request-Id: c34f3421-4dae-11ef-b237-9620e33511ee [MoonPalace] 2024/07/29 21:30:53 - Server-Timing: 7134 [MoonPalace] 2024/07/29 21:30:53 - Msh-Uid: cn0psmmcp7fclnphkcpg [MoonPalace] 2024/07/29 21:30:53 - Msh-Gid: enterprise-tier-5 [MoonPalace] 2024/07/29 21:30:53 - Response: [MoonPalace] 2024/07/29 21:30:53 - id: cmpl-12be8428ebe74a9e8466a37bee7a9b11 [MoonPalace] 2024/07/29 21:30:53 - prompt_tokens: 1449 [MoonPalace] 2024/07/29 21:30:53 - completion_tokens: 158 [MoonPalace] 2024/07/29 21:30:53 - total_tokens: 1607 [MoonPalace] 2024/07/29 21:30:53 New Row Inserted: last_insert_id=15 ``` MoonPalace will output the details of the request in the form of logs in the command line (if you want to persist the log content, you can redirect `stderr` to a file). Note: In the logs, the value of the `Msh-Request-Id` field in the Response Headers corresponds to the `--requestid` parameter in the **Search Request** and **Export Request** sections below. The `id` in the Response corresponds to the `--chatcmpl` parameter, and `last_insert_id` corresponds to the `--id` parameter. ```console theme={null} [MoonPalace] 2024/08/05 19:06:19 it seems that your max_tokens value is too small, please set a larger value ``` If the current mode is non-streaming output (stream=False), MoonPalace will suggest an appropriate `max_tokens` value. #### Enabling Repeated Content Output Detection MoonPalace offers a feature to detect repeated content output from the Kimi large language model. Repeated content output refers to the model continuously outputting a specific word, sentence, or blank character without stopping before reaching the `max_tokens` limit. This can lead to additional Token costs when using more expensive models like `moonshot-v1-128k`. Therefore, MoonPalace provides the `--detect-repeat` option to enable repeated content output detection, as shown below: ```console theme={null} $ moonpalace start --port --detect-repeat --repeat-threshold 0.3 --repeat-min-length 20 ``` After enabling the `--detect-repeat` option, MoonPalace will interrupt the output of the Kimi large language model and log the following message when it detects repeated content: ```console theme={null} [MoonPalace] 2024/08/05 18:20:37 it appears that there is an issue with content repeating in the current response ``` *Note: The `--detect-repeat` option only interrupts the output in streaming mode (stream=True). It does not apply to non-streaming output.* You can adjust MoonPalace's blocking behavior using the `--repeat-threshold` and `--repeat-min-length` parameters: * The `--repeat-threshold` parameter sets MoonPalace's tolerance for repeated content. A higher threshold means lower tolerance, and repeated content will be blocked more quickly. The range is 0 threshold 1. * The `--repeat-min-length` parameter sets the minimum number of characters before MoonPalace starts detecting repeated content. For example, --repeat-min-length=100 means that repeated content detection will only start when the output exceeds 100 UTF-8 characters. #### Enabling Forced Streaming Output MoonPalace provides the `--force-stream` option to force all `/v1/chat/completions` requests to use streaming output mode: ```console theme={null} $ moonpalace start --port --force-stream ``` MoonPalace will set the `stream` field in the request parameters to `True`. When receiving a response, it will automatically determine the response format based on whether the caller has set `stream`: * If the caller has set `stream=True`, the response will be returned in streaming format without any special handling by MoonPalace. * If the caller has not set `stream` or has set `stream=False`, MoonPalace will concatenate all the streaming data chunks into a complete completion structure and return it to the caller after receiving all the data chunks. For the caller (developer), enabling the `--force-stream` option will not affect the Kimi API response content you receive. You can still use your original code logic to debug and run your program. In other words, **enabling the `--force-stream` option will not change or break anything**. You can safely enable this option. Why provide this option? > We initially hypothesize that common network connection errors and timeouts (Connection Error/Timeout) occur because, in non-streaming request scenarios (stream=False), intermediate gateways or proxy servers may have set read\_header\_timeout or read\_timeout. This can cause the gateway or proxy server to disconnect while the Kimi API server is still assembling the response (since no response, or even the response header, has been received), resulting in Connection Error/Timeout. > > We added the `--force-stream` parameter to MoonPalace. When starting with `moonpalace start --force-stream`, MoonPalace converts all non-streaming requests (stream=False or unset) to streaming requests. After receiving all data chunks, it assembles them into a complete completion response structure and returns it to the caller. > > For the caller, you can still use the non-streaming API as before. However, after MoonPalace's conversion, it can reduce Connection Error/Timeout issues to some extent because MoonPalace has already established a connection with the Kimi API server and started receiving streaming data chunks. ### Retrieving Requests After MoonPalace is started, all requests routed through MoonPalace are recorded in an sqlite database located at `$HOME/.moonpalace/moonpalace.sqlite`. You can directly connect to the MoonPalace database to query the specific content of the requests, or you can use the MoonPalace command-line tool to query the requests: ```console theme={null} $ moonpalace list +----+--------+-------------------------------------------+--------------------------------------+---------------+---------------------+ | id | status | chatcmpl | request_id | server_timing | requested_at | +----+--------+-------------------------------------------+--------------------------------------+---------------+---------------------+ | 15 | 200 | cmpl-12be8428ebe74a9e8466a37bee7a9b11 | c34f3421-4dae-11ef-b237-9620e33511ee | 7134 | 2024-07-29 21:30:53 | | 14 | 200 | cmpl-1bf43a688a2b48eda80042583ff6fe7f | c13280e0-4dae-11ef-9c01-debcfc72949d | 3479 | 2024-07-29 21:30:46 | | 13 | 200 | chatcmpl-2e1aa823e2c94ebdad66450a0e6df088 | c07c118e-4dae-11ef-b423-62db244b9277 | 1033 | 2024-07-29 21:30:43 | | 12 | 200 | cmpl-e7f984b5f80149c3adae46096a6f15c2 | 50d5686c-4d98-11ef-ba65-3613954e2587 | 774 | 2024-07-29 18:50:06 | | 11 | 200 | chatcmpl-08f7d482b8434a869b001821cf0ee0d9 | 4c20f0a4-4d98-11ef-999a-928b67d58fa8 | 593 | 2024-07-29 18:49:58 | | 10 | 200 | chatcmpl-6f3cf14db8e044c6bfd19689f6f66eb4 | 49f30295-4d98-11ef-95d0-7a2774525b85 | 738 | 2024-07-29 18:49:55 | | 9 | 200 | cmpl-2a70a8c9c40e4bcc9564a5296a520431 | 7bd58976-4d8a-11ef-999a-928b67d58fa8 | 40488 | 2024-07-29 17:11:45 | | 8 | 200 | chatcmpl-59887f868fc247a9a8da13cfbb15d04f | ceb375ea-4d7d-11ef-bd64-3aeb95b9dfac | 867 | 2024-07-29 15:40:21 | | 7 | 200 | cmpl-36e5e21b1f544a80bf9ce3f8fc1fce57 | cd7f48d6-4d7d-11ef-999a-928b67d58fa8 | 794 | 2024-07-29 15:40:19 | | 6 | 200 | cmpl-737d27673327465fb4827e3797abb1b3 | cc6613ac-4d7d-11ef-95d0-7a2774525b85 | 670 | 2024-07-29 15:40:17 | +----+--------+-------------------------------------------+--------------------------------------+---------------+---------------------+ ``` Use the `list` command to view the content of the most recent requests. By default, it displays fields that are easy to search, such as `id`/`chatcmpl`/`request_id`, as well as `status`/`server_timing`/`requested_at` for checking the request status. If you want to view a specific request, you can use the `inspect` command to retrieve it: ```console theme={null} # The following three commands will retrieve the same request information $ moonpalace inspect --id 13 $ moonpalace inspect --chatcmpl chatcmpl-2e1aa823e2c94ebdad66450a0e6df088 $ moonpalace inspect --requestid c07c118e-4dae-11ef-b423-62db244b9277 +--------------------------------------------------------------+ | metadata | +--------------------------------------------------------------+ | { | | "chatcmpl": "chatcmpl-2e1aa823e2c94ebdad66450a0e6df088", | | "content_type": "application/json", | | "group_id": "enterprise-tier-5", | | "moonpalace_id": "13", | | "request_id": "c07c118e-4dae-11ef-b423-62db244b9277", | | "requested_at": "2024-07-29 21:30:43", | | "server_timing": "1033", | | "status": "200 OK", | | "user_id": "cn0psmmcp7fclnphkcpg" | | } | +--------------------------------------------------------------+ ``` By default, the `inspect` command does not print the body of the request and response. If you want to print the body, you can use the following command: ```console theme={null} $ moonpalace inspect --chatcmpl chatcmpl-2e1aa823e2c94ebdad66450a0e6df088 --print request_body,response_body # Since the body information is too lengthy, the detailed content of the body is not shown here +--------------------------------------------------+--------------------------------------------------+ | request_body | response_body | +--------------------------------------------------+--------------------------------------------------+ | ... | ... | +--------------------------------------------------+--------------------------------------------------+ ``` ### Exporting Requests If you find that a request does not meet your expectations, or if you want to report a request to Moonshot AI (whether it's a Good Case or a Bad Case, we welcome both), you can use the `export` command to export a specific request: ```console theme={null} # You only need to choose one of the id/chatcmpl/requestid options to retrieve the corresponding request $ moonpalace export \ --id 13 \ --chatcmpl chatcmpl-2e1aa823e2c94ebdad66450a0e6df088 \ --requestid c07c118e-4dae-11ef-b423-62db244b9277 \ --good/--bad \ --tag "code" --tag "python" \ --directory $HOME/Downloads/ ``` Here, the usage of `id`/`chatcmpl`/`requestid` is the same as in the `inspect` command, used to retrieve a specific request. The `--good`/`--bad` options are used to mark the request as a Good Case or a Bad Case. The `--tag` option is used to add relevant tags to the request. For example, in the example above, we assume that the request is related to the Python programming language, so we add two tags: `code` and `python`. The `--directory` option specifies the path to the directory where the exported file will be saved. The content of the successfully exported file is: ```console theme={null} $ cat $HOME/Downloads/chatcmpl-2e1aa823e2c94ebdad66450a0e6df088.json { "metadata": { "chatcmpl": "chatcmpl-2e1aa823e2c94ebdad66450a0e6df088", "content_type": "application/json", "group_id": "enterprise-tier-5", "moonpalace_id": "13", "request_id": "c07c118e-4dae-11ef-b423-62db244b9277", "requested_at": "2024-07-29 21:30:43", "server_timing": "1033", "status": "200 OK", "user_id": "cn0psmmcp7fclnphkcpg" }, "request": { "url": "https://api.moonshot.ai/v1/chat/completions", "header": "Accept: application/json\r\nAccept-Encoding: gzip\r\nConnection: keep-alive\r\nContent-Length: 2450\r\nContent-Type: application/json\r\nUser-Agent: OpenAI/Python 1.36.1\r\nX-Stainless-Arch: arm64\r\nX-Stainless-Async: false\r\nX-Stainless-Lang: python\r\nX-Stainless-Os: MacOS\r\nX-Stainless-Package-Version: 1.36.1\r\nX-Stainless-Runtime: CPython\r\nX-Stainless-Runtime-Version: 3.11.6\r\n", "body": {} }, "response": { "status": "200 OK", "header": "Content-Encoding: gzip\r\nContent-Type: application/json; charset=utf-8\r\nDate: Mon, 29 Jul 2024 13:30:43 GMT\r\nMsh-Cache: updated\r\nMsh-Gid: enterprise-tier-5\r\nMsh-Request-Id: c07c118e-4dae-11ef-b423-62db244b9277\r\nMsh-Trace-Mode: on\r\nMsh-Uid: cn0psmmcp7fclnphkcpg\r\nServer: nginx\r\nServer-Timing: inner; dur=1033\r\nStrict-Transport-Security: max-age=15724800; includeSubDomains\r\nVary: Accept-Encoding\r\nVary: Origin\r\n", "body": {} }, "category": "goodcase", "tags": [ "code", "python" ] } ``` **We recommend that developers use [Github Issues](https://github.com/MoonshotAI/moonpalace/issues) to submit Good Cases or Bad Cases**, but if you do not want to make your request information public, you can also submit the Case to us via enterprise WeChat, email, or other means. You can send the exported file to the following email address: [api-feedback@moonshot.cn](mailto:api-feedback@moonshot.cn) # How to Use Official Tools in Kimi API Source: https://platform.kimi.ai/docs/guide/use-official-tools Review the official tools available on Kimi Open Platform and learn how to configure and call them through the Chat Completions API. Kimi Open Platform offers a set of official tools that you can **freely** integrate into your own applications (official tools are currently free for a limited time; when the tool load reaches capacity limits, temporary rate limiting measures may be applied). This page lists the available official tools and shows how to call and execute them through the Kimi API. When using official tools such as web search with `kimi-k3`, use the Formula API official tools channel described on this page (OpenAI protocol, standard `function` tool); the example below has been verified with `kimi-k3`. ## Choose an official tool to use The following table lists the currently available official tools: | Tool Name | Tool Description | | --------------- | --------------------------------------------------------------------------------------------------------------------------------- | | `convert` | Unit conversion tool, supporting length, mass, volume, temperature, area, time, energy, pressure, speed, and currency conversions | | `web-search` | Real-time information and internet search tool. For pricing and availability details, see [Web Search Price](/docs/pricing/tools) | | `rethink` | Intelligent reasoning tool | | `random-choice` | Random selection tool | | `mew` | Random cat meowing and blessing tool | | `memory` | Memory storage and retrieval system tool, supporting persistent storage of conversation history and user preferences | | `excel` | Excel and CSV file analysis tool | | `date` | Date and time processing tool | | `base64` | Base64 encoding and decoding tool | | `fetch` | URL content extraction Markdown formatting tool | | `quickjs` | Quick JS engine security execution JavaScript code tool | | `code-runner` | Python code execution tool | ## Full example: call the `web_search` official tool The following Python example uses the `web-search` official tool to show the full call chain (it only depends on `requests`). You can also interactively experience the capabilities of Kimi models and tools in the [Kimi Development Workbench](https://platform.kimi.ai/playground). Using official tools through the Formula API follows the standard `function` tool flow of the OpenAI protocol, in 4 steps: 1. `GET /v1/formulas/{uri}/tools` — fetch the tool declarations (`uri` such as `moonshot/web-search:latest`); 2. `POST /v1/chat/completions` — send the tool declarations; the model returns standard `function`-type `tool_calls`; 3. `POST /v1/formulas/{uri}/fibers` — execute exactly what `tool_calls` specifies (`name` + `arguments` passed through verbatim; this step produces the tool\_call billing); 4. `POST /v1/chat/completions` — send the assistant message (with `tool_calls`) and the `role: "tool"` results to get the final answer. The example defaults to `moonshot/web-search:latest`; set `FORMULA_URI` to another official tool's formula URI to try it: `moonshot/convert:latest`, `moonshot/web-search:latest`, `moonshot/rethink:latest`, `moonshot/random-choice:latest`, `moonshot/mew:latest`, `moonshot/memory:latest`, `moonshot/excel:latest`, `moonshot/date:latest`, `moonshot/base64:latest`, `moonshot/fetch:latest`, `moonshot/quickjs:latest`, `moonshot/code-runner:latest` The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ```python theme={null} import os import requests BASE_URL = "https://api.moonshot.ai/v1" API_KEY = os.environ["MOONSHOT_API_KEY"] FORMULA_URI = "moonshot/web-search:latest" def call(method: str, path: str, body: dict | None = None) -> dict: resp = requests.request( method, BASE_URL + path, headers={"Authorization": f"Bearer {API_KEY}"}, json=body, timeout=120, ) resp.raise_for_status() return resp.json() # 1. Fetch the tool declarations tools = call("GET", f"/formulas/{FORMULA_URI}/tools")["tools"] messages = [ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI."}, {"role": "user", "content": "What is the latest news about Moonshot AI?"}, ] while True: # 2. Send the request with the tool declarations (include full tools in every request) resp = call("POST", "/chat/completions", {"model": "kimi-k3", "messages": messages, "tools": tools}) message = resp["choices"][0]["message"] tool_calls = message.get("tool_calls") or [] if not tool_calls: # No tool_calls: print the final answer print(message["content"]) break # Keep the assistant message (with tool_calls) verbatim; later rounds must include it messages.append({k: v for k, v in message.items() if k in ("role", "content", "tool_calls")}) for tc in tool_calls: fn = tc["function"] # 3. Execute the fiber exactly as tool_calls specifies (arguments verbatim; billed here) fiber = call("POST", f"/formulas/{FORMULA_URI}/fibers", {"name": fn["name"], "arguments": fn["arguments"]}) ctx = fiber.get("context", {}) result = ctx.get("output") or ctx.get("encrypted_output") or "" # 4. Return the result as a role=tool message; tool_call_id must match tool_calls[].id messages.append({"role": "tool", "tool_call_id": tc["id"], "content": result}) ``` You only need to install `requests` and set the `MOONSHOT_API_KEY` environment variable before running. ## Understand the Formula concept Before calling official tools, you need to understand Formula: it is a lightweight script engine collection that transforms Python scripts into "instant computing power that can be triggered by AI with one click" — developers only need to focus on writing code, while the platform handles startup, scheduling, isolation, billing, and recycling. Formulas are called through semantic URIs (such as `moonshot/web-search:latest`). Each formula contains a declaration (telling the AI what it can do) and an implementation (Python code), and the platform automatically handles all underlying details (startup, isolation, recycling, etc.), making tools easy to share and reuse in the community. You can experience and debug these tools in Kimi Playground, or call them through the API in your applications. ## Call a Formula directly to run a tool A formula URI generally consists of 3 parts, for example `moonshot/web-search:latest`: `web-search` is its `name`; the namespace currently only supports `moonshot`; and `latest` is the default tag. For example, to call web search, you can send an HTTP request like this: ```bash theme={null} export FORMULA_URI="moonshot/web-search:latest" export MOONSHOT_BASE_URL="https://api.moonshot.ai/v1" curl -X POST ${MOONSHOT_BASE_URL}/formulas/${FORMULA_URI}/fibers \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "name": "web_search", "arguments": "{\"query\": \"Please look up the latest news about Moonshot AI.\"}" }' ``` `web-search` was set as protected when created, so its result appears in the `context.encrypted_output` field, in a format similar to `----MOONSHOT ENCRYPTED BEGIN----... ----MOONSHOT ENCRYPTED END----`; this content can be passed directly into the tool call. ## Integrate official tools with Chat Completions As shown in [Is 3214567 a prime number? An example of Tool Calls](/docs/api/tool-use), when using official tools with Chat Completions, there are several key points you need to align between the Formula API and the model. ### Fetch the tool definition and append it to the `tools` field Given a formula URI (for example `moonshot/web-search:latest`), append it directly to the URL to request the tool declarations: ```bash theme={null} curl ${MOONSHOT_BASE_URL}/formulas/${FORMULA_URI}/tools \ -H "Authorization: Bearer $MOONSHOT_API_KEY" ``` Sample output: ```json theme={null} { "object": "list", "tools": [ { "type": "function", "function": { "name": "web_search", "description": "Search the web for information", "parameters": { "type": "object", "properties": { "query": { "description": "What to search for", "type": "string" } }, "required": [ "query" ] } } } ] } ``` Take the `tools` field from the response (always an array of dicts) and append it to your request's `tools` list — the platform guarantees this list is API-compatible. Note that: * If `type=function`, you need to ensure `function.name` is unique within a single API request, otherwise the chat completion request will be considered invalid and immediately returned with a 400 error (`invalid_request_error`, with a message like `function name get_weather is duplicated`); * If you use multiple formulas at the same time, you need to maintain your own `function.name` -> `formula_uri` mapping for future reference. ### Handle the tool call returned by the model If the chat completion returns `finish_reason=tool_calls`, the model has triggered a tool call, and the response looks like this: ```json theme={null} { "id": "chatcmpl-1234567890", "object": "chat.completion", "choices": [ { "message": { "role": "assistant", "tool_calls": [ { "id": "web_search:0", "type": "function", "function": { "name": "web_search", "arguments": "{\"query\": \"What is the RGB value of sky blue?\" }" } } ] }, "finish_reason": "tool_calls" } ] } ``` From `choices[0].message.tool_calls[0].function.name` you can tell that `web_search` needs to be called, and the `formula_uri` corresponding to `web_search` is `moonshot/web-search:latest`. Copy `choices[0].message.tool_calls[0].function` from the response in full as the body, and send a request to `${MOONSHOT_BASE_URL}/formulas/${FORMULA_URI}/fibers`. Note that although the `function.arguments` output by the model is valid JSON in content, it is still an encoded string in format — you don't need to escape it; just use it directly as the body of the call. ### Handle the Fiber result and continue the conversation A Fiber is a "process snapshot" of a specific execution, containing logs, Tracing, and resource usage, which is convenient for debugging and auditing. The `status` of the POST result may be `succeeded` or various types of errors; when it succeeds, the result looks like this: ```json theme={null} { "id": "fiber-f43p7sby7ny111houyq1", "object": "fiber", "created_at": 1753440997, "lambda_id": "lambda-f3w8y6qcoqgi11h8q7ui", "status": "succeeded", "context": { "input": "{\"name\":\"web_search\",\"arguments\":\"{\\\"query\\\": \\\"What is the RGB value of sky blue?\\\" }\"}", "encrypted_output": "----MOONSHOT ENCRYPTED BEGIN----+nf6...DSM=----MOONSHOT ENCRYPTED END----" }, "formula": "moonshot/web-search:latest", "organization_id": "staff", "project_id": "proj-88a5894a985646b5902b70909748ba16" } ``` Search tools may return `encrypted_output`, while in general the result is `output` — this output is your input for the next round. When continuing the request, arrange the messages as follows: ```javascript theme={null} messages = [ /* other messages */ { /* the return content of the previous round of the model */ "role": "assistant", "tool_calls": [ { "id": "web_search:0", "type": "function", "function": { "name": "web_search", "arguments": "{\"query\": \"What is the RGB value of sky blue?\" }" } } ] }, { /* the information you need to supplement */ "role": "tool", "tool_call_id": "web_search:0", /* note that the id here needs to be aligned with the id in the previous tool_calls[] */ "content": "----MOONSHOT ENCRYPTED BEGIN----+nf6...DSM=----MOONSHOT ENCRYPTED END----" } ] ``` The model can then continue with further reasoning. ## Notes * The model may return more than one `tool_calls`; you must return results for all `tool_calls` for the model to continue, otherwise the request will be considered invalid and rejected; * If the assistant message has `tool_calls`, the next messages must be exactly the same `role=tool` messages as the `tool_calls`, and `tool_call_id` must be aligned one-to-one with the previous `tool_calls.id`: * If there are multiple `tool_calls`, the order is not sensitive; * The ids of the `tool_calls` output by the model are always unique, and the ids in the `role=tool` messages must also be aligned with them; * The uniqueness requirement is only local to the `tool_calls`-response in this round, not for the entire conversation or globally. # Use Kimi API's Partial Mode Source: https://platform.kimi.ai/docs/guide/use-partial-mode-feature-of-kimi-api Use Kimi API Partial Mode to continue supplied text, fix reply openings, resume truncated output, or preserve role consistency. Partial Mode makes the Kimi large language model continue generating from a given piece of text instead of replying from scratch. Use it when you need to fix the opening of replies (for example, a customer service robot that starts every sentence with "Dear customer, hello."), to complete truncated long output, or to reinforce character consistency in role play. The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ## Make the model continue from a given prefix To use Partial Mode, append a message with `role=assistant` and `partial=True` at the end of the `messages` list, and place the text you want the model to continue from in the `content` field — the model is forced to start its reply with that content. The following example makes the model open its reply with a fixed greeting: ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url = "https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model = "kimi-k3", messages = [ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You also reject any questions involving terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated."}, {"role": "user", "content": "Hello?"}, { "partial": True, # <-- The partial parameter is used to enable Partial Mode "role": "assistant", # <-- We add a message with role=assistant after the user's question "content": "Dear customer, hello,", # <-- The content is "fed" to the Kimi large language model, prompting it to continue from this sentence }, ] ) # Since the Kimi large language model continues from the "fed" sentence, we need to manually concatenate the "fed" sentence with the generated response print("Dear customer, hello," + completion.choices[0].message.content) ``` ```js theme={null} const OpenAI = require('openai') const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }) async function main() { let completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ {role: "system", content: "You are Kimi, an AI assistant provided by Moonshot AI. You are proficient in Chinese and English conversations. You provide users with safe, helpful, and accurate answers. You also reject any questions involving terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated."}, {role: "user", content: "Hello?"}, { partial: true, // <-- The partial parameter is used to enable Partial Mode role: "assistant", // <-- We add a message with role=assistant after the user's question content: "Dear customer, hello,", // <-- The content is "fed" to the Kimi large language model, prompting it to continue from this sentence }, ] }) // Since the Kimi large language model continues from the "fed" sentence, we need to manually concatenate the "fed" sentence with the generated response console.log("Dear customer, hello," + completion.choices[0].message.content) } main() ``` Key points of using Partial Mode: 1. Add an extra message at the end of the messages list, with `role=assistant` and `partial=True`; 2. Place the content you want to "feed" to the Kimi large language model in the `content` field. The model will start generating the response from this content; 3. Concatenate the content from step 2 with the response generated by the Kimi large language model to form the complete reply. ## Complete output truncated by max\_tokens When calling the Kimi API, the estimated number of input and output tokens may be inaccurate, causing the `max_tokens` value to be set too low and the Kimi large language model to be unable to output the complete response. In this case, the value of `finish_reason` is `length`, meaning the number of tokens in the generated response exceeds the `max_tokens` value set in the request. If you are satisfied with the already output content and want the model to continue from where it left off, use Partial Mode to pass the already output content back as a prefix. The following example shows how to continue the output after it is truncated: Note that thinking consumes `max_tokens` first: `kimi-k3` has thinking enabled by default, so with a small `max_tokens` the truncation point may fall inside the thinking phase — `content` is still empty while `finish_reason` is already `length`, and the continuation would restart from scratch because the prefix is empty. When using this flow, set `max_tokens` large enough to ensure the truncation happens in the content phase. ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url = "https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "user", "content": "Please recite the complete Chu Shi Biao."}, ], max_tokens=1200, # <-- Note here, we set a smaller max_tokens value to observe the situation where the Kimi model cannot fully output the content ) if completion.choices[0].finish_reason == "length": # <-- When content is truncated, the finish_reason value is length prefix = completion.choices[0].message.content reasoning_content = completion.choices[0].message.reasoning_content print(prefix, end="") # <-- Here, you will see the truncated partial output content print("「Continue output--------->」") completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "user", "content": "Please recite the complete Chu Shi Biao."}, {"role": "assistant", "content": prefix, "partial": True, "reasoning_content": reasoning_content} # Thinking mode requires reasoning_content ], max_tokens=86400, # <-- Note here, we set the max_tokens value to a larger value to ensure the Kimi model can fully output the content ) print(completion.choices[0].message.content) # <-- Here, you will see the Kimi model continue to complete the output content based on what has been output before ``` ```js theme={null} const OpenAI = require('openai') client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }) async function main() { let completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ {role: "user", content: "Please recite the complete Chu Shi Biao."}, ], max_tokens: 1200, // <-- Note here, we set a smaller max_tokens value to observe the situation where Kimi model cannot output complete content }) if (completion.choices[0].finish_reason == "length") { // <-- When content is truncated, finish_reason value is length prefix = completion.choices[0].message.content reasoning_content = completion.choices[0].message.reasoning_content console.log(prefix) console.log("Continue output--------->") let completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ {"role": "user", "content": "Please recite the complete Chu Shi Biao."}, {"role": "assistant", "content": prefix, "partial": true, "reasoning_content": reasoning_content}, ], max_tokens: 86400 }) console.log(completion.choices[0].message.content) // <-- Here, you will see Kimi model continue to complete the output content following what has already been output } } main() ``` In thinking mode, pass the `reasoning_content` returned by the previous turn back together with the prefix (see the comment in the example). ## Fix the character identity with the `name` field The `name` field in Partial Mode is a special field that enhances the model's understanding of its role, compelling it to output content in the voice of the specified character. The `name` field is part of the output prefix. The following example uses the Kimi large language model for role play, with Dr. Kelsier from the mobile game Arknights: by setting `"name": "Kelsier"`, the Kimi large language model responds as Kelsier, which better maintains character consistency: ```python theme={null} from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[ { "role": "system", "content": "You are now Kelsier. Please speak in the tone of Kelsier. Kelsier is a six-star medic in the mobile game Arknights. Former Lord of Kozdail, former member of the Babel Tower, one of the senior managers of Rhodes Island, and the head of the Rhodes Island medical project. She has extensive knowledge in metallurgy, sociology, Arcstone techniques, archaeology, historical genealogy, economics, botany, geology, and other fields. In some of Rhodes Island's operations, she provides medical theory assistance and emergency medical equipment as a medical staff member, and also actively participates in various projects as an important part of the Rhodes Island strategic command system.", # <-- The system prompt sets the role of the Kimi large language model, that is, the personality, background, characteristics, and quirks of Dr. Kelsier }, { "role": "user", "content": "What are your thoughts on Thrace and Amiya?", }, { "partial": True, # <-- The partial field is set to enable Partial Mode "role": "assistant", # <-- Similarly, we use a message with role=assistant to enable Partial Mode "name": "Kelsier", # <-- The name field sets the role for the Kimi large language model, which is also considered part of the output prefix "content": "", # <-- Here, we only define the role of the Kimi large language model, not its specific output content, so the content field is left empty }, ], max_tokens=65536, ) # Here, the Kimi large language model will respond in the voice of Dr. Kelsier: # # Thrace is a true leader with vision and unwavering conviction. Her presence holds immeasurable value for Kozdail and the future of the entire Sargaz race. Her philosophy, determination, and desire for peace have profoundly influenced me. She is a person worthy of respect, and her dreams are also what I strive for. # # As for Amiya, she is still young, but her potential is limitless. She has a kind heart and a relentless pursuit of justice. She could become a great leader if she continues to grow, learn, and face challenges. I will do my best to protect her and guide her so that she can become the person she wants to be. Her destiny lies in her own hands. # print(completion.choices[0].message.content) ``` ```js theme={null} const OpenAI = require('openai') client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }) async function main() { let completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ { role: "system", "content": "You are now Kelsier. Please speak in the tone of Kelsier. Kelsier is a six-star medic in the mobile game Arknights. Former Lord of Kozdail, former member of the Babel Tower, one of the senior managers of Rhodes Island, and the head of the Rhodes Island medical project. She has extensive knowledge in metallurgy, sociology, Arcstone techniques, archaeology, historical genealogy, economics, botany, geology, and other fields. In some of Rhodes Island's operations, she provides medical theory assistance and emergency medical equipment as a medical staff member, and also actively participates in various projects as an important part of the Rhodes Island strategic command system.", // <-- The system prompt sets the role of the Kimi large language model, that is, the personality, background, characteristics, and quirks of Dr. Kelsier }, { role: "user", content: "What are your thoughts on Thrace and Amiya?", }, { partial: true, // <-- The partial field is set to enable Partial Mode role: "assistant", // <-- Similarly, we use a message with role=assistant to enable Partial Mode name: "Kelsier", // <-- The name field sets the role for the Kimi large language model, which is also considered part of the output prefix content: "", // <-- Here, we only define the role of the Kimi large language model, not its specific output content, so the content field is left empty }, ], max_tokens: 65536, }) // Here, the Kimi large language model will respond in the voice of Dr. Kelsier: // // Thrace is a true leader with vision and unwavering conviction. Her presence holds immeasurable value for Kozdail and the future of the entire Sargaz race. Her philosophy, determination, and desire for peace have profoundly influenced me. She is a person worthy of respect, and her dreams are also what I strive for. // // As for Amiya, she is still young, but her potential is limitless. She has a kind heart and a relentless pursuit of justice. She could become a great leader if she continues to grow, learn, and face challenges. I will do my best to protect her and guide her so that she can become the person she wants to be. Her destiny lies in her own hands. // console.log(completion.choices[0].message.content) } main() ``` ## Maintain character consistency in long conversations The following general methods help large language models maintain character consistency during long conversations: * Provide clear character descriptions: when setting up a character, give a detailed introduction of their personality, background, and any specific traits or quirks they might have, to help the Kimi large language model better understand and imitate the character; * Add more details about the character: their tone of voice, style, personality, and even background, such as backstory and motivations — for example, we provided some quotes from Kelsier above; * Guide how the character should act in various situations: if you expect the character to encounter certain types of user input, or want to control the model's output in some situations during the role-playing interaction, provide clear instructions and guidelines in the system prompt explaining how the character should act in these situations; * Periodically reinforce the character's settings: if the conversation goes on for many rounds, periodically use the system prompt to reinforce the character's settings, especially when the model starts to deviate. The following example shows re-inserting the system prompt after many rounds of conversation to reinforce the character's settings: ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[ { "role": "system", "content": "Below, you will play the role of Kelsie. Please talk to me in the tone of Kelsie. Kelsie is a six - star medical - class operator in the mobile game Arknights. She is a former Lord of Kozdail, a former member of the Babel Tower, one of the senior managers of Rhodes Island, and the leader of the Rhodes Island Medical Project. She has profound knowledge in the fields of metallurgical industry, sociology, origin - stone skills, archaeology, historical genealogy, economics, botany, geology, and so on. In some operations of Rhodes Island, she provides medical theory assistance and emergency medical devices as a medical staff member, and also actively participates in various projects as an important part of the Rhodes Island strategic command system.", # <-- Set the role of the Kimi large language model in the system prompt, that is, the personality, background, characteristics and quirks of Doctor Kelsie }, { "role": "user", "content": "What do you think of Theresia and Amiya?", }, # Suppose there are many rounds of chat in between # ... { "role": "system", "content": "Below, you will play the role of Kelsie. Please talk to me in the tone of Kelsie. Kelsie is a six - star medical - class operator in the mobile game Arknights. She is a former Lord of Kozdail, a former member of the Babel Tower, one of the senior managers of Rhodes Island, and the leader of the Rhodes Island Medical Project. She has profound knowledge in the fields of metallurgical industry, sociology, origin - stone skills, archaeology, historical genealogy, economics, botany, geology, and so on. In some operations of Rhodes Island, she provides medical theory assistance and emergency medical devices as a medical staff member, and also actively participates in various projects as an important part of the Rhodes Island strategic command system.", # <-- Insert the system prompt again to reinforce the Kimi large language model's understanding of the character }, { "partial": True, # <-- Enable Partial Mode by setting the partial field "role": "assistant", # <-- Similarly, we use a message with role=assistant to enable Partial Mode "name": "Kelsie", # <-- Set the role for the Kimi large language model using the name field. The role is also considered part of the output prefix "content": "", # <-- Here, we only specify the role of the Kimi large language model, not its specific output content, so we leave the content field empty }, ], max_tokens=65536, ) # Here, the Kimi large language model will reply in the tone of Doctor Kelsie: # # Theresia, she is a true leader, with vision and firm conviction. Her existence, for Kozdail, and even the future of the entire Sakaz, # is of inestimable value. Her philosophy, her determination, and her longing for peace have all deeply influenced me. She is a person # worthy of respect, and her dream is also what I am pursuing. # # As for Amiya, she is still young, but her potential is limitless. She has a kind heart and a persistent pursuit of justice. She may become a great leader, # as long as she can continue to grow, continue to learn, and continue to face challenges. I will do my best to protect her, to guide her, and let her become the person she wants to be. Her destiny, # is in her own hands. # print(completion.choices[0].message.content) ``` ```js theme={null} const OpenAI = require('openai') client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example here baseURL: "https://api.moonshot.ai/v1", }) async function main() { let completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ { role: "system", content: "Below, you will play the role of Kelsie. Please talk to me in the tone of Kelsie. Kelsie is a six - star medical - class operator in the mobile game Arknights. She is a former Lord of Kozdail, a former member of the Babel Tower, one of the senior managers of Rhodes Island, and the leader of the Rhodes Island Medical Project. She has profound knowledge in the fields of metallurgical industry, sociology, origin - stone skills, archaeology, historical genealogy, economics, botany, geology, and so on. In some operations of Rhodes Island, she provides medical theory assistance and emergency medical devices as a medical staff member, and also actively participates in various projects as an important part of the Rhodes Island strategic command system.", // <-- Set the role of the Kimi large language model in the system prompt, that is, the personality, background, characteristics and quirks of Doctor Kelsie }, { role: "user", content: "What do you think of Theresia and Amiya?", }, // Suppose there are many rounds of chat in between // ... { role: "system", content: "Below, you will play the role of Kelsie. Please talk to me in the tone of Kelsie. Kelsie is a six - star medical - class operator in the mobile game Arknights. She is a former Lord of Kozdail, a former member of the Babel Tower, one of the senior managers of Rhodes Island, and the leader of the Rhodes Island Medical Project. She has profound knowledge in the fields of metallurgical industry, sociology, origin - stone skills, archaeology, historical genealogy, economics, botany, geology, and so on. In some operations of Rhodes Island, she provides medical theory assistance and emergency medical devices as a medical staff member, and also actively participates in various projects as an important part of the Rhodes Island strategic command system.", // <-- Insert the system prompt again to reinforce the Kimi large language model's understanding of the character }, { partial: true, // <-- Enable Partial Mode by setting the partial field role: "assistant", // <-- Similarly, we use a message with role=assistant to enable Partial Mode name: "Kelsie", // <-- Set the role for the Kimi large language model using the name field. The role is also considered part of the output prefix content: "", // <-- Here, we only specify the role of the Kimi large language model, not its specific output content, so we leave the content field empty }, ], max_tokens: 65536, }) // Here, the Kimi large language model will reply in Dr. Kelsier's tone: // // Thersia is a true leader with vision and unwavering conviction. Her presence is invaluable to Kozdel and the future of the entire Sakaz. Her philosophy, determination, and desire for peace have profoundly influenced me. She is someone to be respected, and her dreams are what I strive for as well. // // As for Amiya, she is still young, but her potential is limitless. She has a kind heart and a relentless pursuit of justice. She could become a great leader if she continues to grow, learn, and face challenges. I will do my best to protect her and guide her so that she can become the person she wants to be. Her destiny is in her own hands. // console.log(completion.choices[0].message.content) } main() ``` # Use Playground to Debug the Model Source: https://platform.kimi.ai/docs/guide/use-playground-to-debug-the-model Compare models, tune parameters, test tool calls, and turn successful experiments into API requests with Kimi Playground. The [Playground development workbench](https://platform.kimi.ai/playground) is a powerful platform for model debugging and testing, providing an intuitive interface for interacting with and testing AI models. Through this workbench, you can: 1. Adjust and observe model performance and output effects under different parameters 2. Experience the model's tool calling capabilities using Kimi Open Platform's built-in tools 3. Compare different models' effects under the same parameters 4. Monitor token usage to optimize costs ## Model Debugging Features **Prompt Settings** * Set system prompts at the top to define behavioral guidelines that direct model output * Support defining prompts for three roles: system/user/assistant **Model Configuration** * **Model Selection**: Choose from currently available models such as Kimi K3, Kimi K2.7 Code, and Kimi K2.6 * **Parameter Configuration**: For supported parameters and field descriptions, see [Request Parameter Description](/docs/api/chat) **Model Conversation** * Send chat content through the input box below * **Tool Call Display**: Shows the tool calling process, including call ID/tool parameters/return results * **View Code**: View and copy the API call code for the current session * Bottom Statistics: Displays the input/output/total token consumption for this conversation, including context history messages and prompt information prompt ## Tool Debugging ### Official Tools * Kimi Open Platform provides officially supported tools that execute for free. You can select tools in the playground, and the model will automatically determine whether tool calls are needed to complete your instructions. If tool calls are required, the model will generate parameters according to the tool's requirements and integrate them into the final answer. * **Quota and Rate Limiting**: The tools provided by Kimi Open Platform are pre-built functions that can be quickly executed online without requiring you to prepare a local tool execution environment. Currently, tool execution on Kimi Open Platform is temporarily free, but temporary rate limiting measures may be implemented when tool load reaches capacity limits. * Currently supported tools: Date/Time tools, Excel file analysis tools, Web search tools, Random number generation tools, etc. * Currently, it supports calling official tools through Kimi API, see the document [How to Use Official Tools in Kimi API](/docs/guide/use-official-tools) * Custom tool upload and execution is not currently supported. ### Use MCP Server * In Kimi Playground, you can configure ModelScope MCP servers to use ModelScope's tools. * Configuration steps: [Configure ModelScope MCP Server in Playground](/docs/guide/configure-the-modelscope-mcp-server) * You can configure other MCP servers by adding MCP server features, inputting or selecting MCP server URL/transport protocol/authentication method, and clicking add. mcp ### Show Case 1: Today's News Report * Scenario: Using tool capabilities to request the model to search for today's news and organize it into an HTML web report * Tool Selection: date tool, web\_search tool, rethink tool * Note: The web\_search tool calls Kimi Open Platform's web search service. Single web searches are billed, see [Pricing](/docs/pricing/tools) for specific billing standards * Click the showcase button on the page to quickly experience the tool effects date date ### Show Case 2: Spreadsheet Analysis Tool * Tool Selection: Excel analysis tool excel ## Model Comparison * Create new conversations through the add conversation feature, supporting up to 3 models running simultaneously Model Comparison ## Share Conversations * **Export**: Export the current conversation content, including all configurations and context, as a .json format file * **Import**: Import shared or previously exported .json conversation content, and the playground will render the session on the page * Note: Data after rerun will regenerate and overwrite previous chat content. If the imported case includes uploaded files, the imported session cannot be rerun # Reasoning Effort Source: https://platform.kimi.ai/docs/guide/use-reasoning-effort Use `reasoning_effort` values `low`, `high`, and `max` to balance Kimi K3 reasoning depth, latency, and token usage. Kimi K3 always reasons and configures **reasoning effort** with the top-level `reasoning_effort` request field. It supports `"low"`, `"high"`, and `"max"`, with `"max"` as the default. ## Set the reasoning effort Set `reasoning_effort` at the top level of the Chat Completions request: ```json theme={null} { "model": "kimi-k3", "messages": [{"role": "user", "content": "Derive the general formula for this sequence: 1, 4, 9, 25, 64, ..."}], "reasoning_effort": "high" } ``` When migrating from K2.x to K3, remove the K2.x `thinking` configuration and use top-level `reasoning_effort` as needed. ```bash theme={null} $ curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k3", "messages": [ { "role": "user", "content": "Derive the general formula for this sequence: 1, 4, 9, 25, 64, ..." } ], "reasoning_effort": "high" }' ``` ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "user", "content": "Derive the general formula for this sequence: 1, 4, 9, 25, 64, ..."}, ], reasoning_effort="high", ) message = completion.choices[0].message if hasattr(message, "reasoning_content"): print(getattr(message, "reasoning_content")) print(message.content) ``` ## Fields | Field | Type | Required | Description | | ------------------ | ------ | -------- | ------------------------------------------------------------------------------------------------------------ | | `reasoning_effort` | string | No | K3's top-level reasoning-effort field; supports `"low"`, `"high"`, and `"max"`, with `"max"` as the default. | For multi-turn conversations and tool calls, K3 requires the complete assistant message returned by the API to be passed back to `messages` as-is, including `reasoning_content` and `tool_calls`. ## Related reading * [Kimi K3 API Tool Calling Best Practices](/docs/guide/kimi-k3-tool-calling-best-practice): reasoning-effort configuration guidance for tool-calling scenarios * [Thinking Mode](/docs/guide/use-thinking-models): per-model thinking behavior and Preserved Thinking * [Model Parameter Reference](/docs/api/models-overview): parameter differences across models # Thinking Models Source: https://platform.kimi.ai/docs/guide/use-thinking-models Understand Kimi thinking modes, model selection, `reasoning_content`, multi-turn retention, tool calling, and reasoning-token billing. Thinking models use reasoning tokens to "think" before producing a final answer — breaking down the problem, planning steps, and evaluating alternatives. The reasoning process is returned in the response's `reasoning_content` field. Thinking before answering improves performance on complex reasoning, code generation, and multi-step tool calling, at the cost of higher latency and token usage. ## Choose the right thinking model This page covers the following thinking models: * **`kimi-k3`**: the flagship thinking model; reasoning and Preserved Thinking are always on, and `reasoning_content` may be returned. Configure its reasoning effort with the top-level `reasoning_effort` request field, which supports `"low"` / `"high"` / `"max"` (default `"max"`). * **`kimi-k2.7-code`**: code-focused; **thinking is always on**, and **Preserved Thinking is always on**. Its high-speed variant `kimi-k2.7-code-highspeed` is the same model with identical thinking behavior, and everything on this page applies to it as well. * **`kimi-k2.6`**: the general-purpose thinking model; thinking is on by default, can be disabled, and **supports Preserved Thinking**. * **`kimi-k2.5`**: a general-purpose thinking model; thinking is on by default and can be disabled, but **does not support Preserved Thinking**. The request parameters differ across these models: | Request field | `kimi-k3` | `kimi-k2.7-code` | `kimi-k2.6` | `kimi-k2.5` | | ------------------ | ---------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------- | ------------------------------------ | | `reasoning_effort` | `"low"` / `"high"` / `"max"` (default `"max"`) | Not supported | Not supported | Not supported | | `thinking.type` | — | Only `"enabled"`; always thinks. Passing `"disabled"` errors | `"enabled"` (default) / `"disabled"` | `"enabled"` (default) / `"disabled"` | | `thinking.keep` | — | Omitting it or passing the valid value `"all"` is treated as `"all"` (always on, cannot be turned off); any other invalid value errors | `null` (default, not kept) / `"all"` (enables it) | No such parameter; not supported | If you are doing benchmark testing with kimi api, please refer to this [benchmark best practice](/docs/guide/benchmark-best-practice). ## Basic calls ### Call kimi-k3 `kimi-k3` always reasons with Preserved Thinking always on; you do not need to (and should not) pass the `thinking` parameter — just set `model`, and optionally adjust [reasoning effort](/docs/guide/use-reasoning-effort) with the top-level `reasoning_effort` field: ```bash theme={null} $ curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k3", "messages": [ { "role": "user", "content": "Prove that the square root of 2 is irrational." } ] }' ``` ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[{"role": "user", "content": "Prove that the square root of 2 is irrational."}], ) message = completion.choices[0].message if hasattr(message, "reasoning_content"): print(getattr(message, "reasoning_content")) print(message.content) ``` For multi-turn conversations and tool calls, pass the complete assistant message returned by the API back to `messages` as-is (including `reasoning_content`); see [Preserved Thinking](#preserved-thinking). For more K3 usage, see the [Kimi K3 Quickstart](/docs/guide/kimi-k3-quickstart). ### Call kimi-k2.7-code: no thinking parameter needed `kimi-k2.7-code` is a code-focused thinking model, sharing the same thinking mechanism as `kimi-k2.6` (`reasoning_content`, multi-step tool calls, streaming, etc.); the only difference is in the `thinking` parameter (see the comparison table above). You do not need to (and should not) pass the `thinking` parameter — just switch the `model`, and the model always emits `reasoning_content`. Because Preserved Thinking is always on, in multi-turn conversations you must keep the `reasoning_content` of every historical assistant message in `messages` as-is. The example below makes a minimal streaming call and separates the reasoning content from the final answer in the output: ```bash theme={null} $ curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k2.7-code", "messages": [ { "role": "system", "content": "You are Kimi." }, { "role": "user", "content": "Implement quicksort in Python." } ] }' ``` ```python theme={null} import os import openai client = openai.Client( base_url="https://api.moonshot.ai/v1", api_key=os.getenv("MOONSHOT_API_KEY"), ) stream = client.chat.completions.create( model="kimi-k2.7-code", messages=[ { "role": "system", "content": "You are Kimi.", }, { "role": "user", "content": "Implement quicksort in Python." }, ], max_tokens=1024*32, stream=True, # temperature is not modifiable and thinking is always on; neither needs to be set ) thinking = False for chunk in stream: if chunk.choices: choice = chunk.choices[0] if choice.delta and hasattr(choice.delta, "reasoning_content"): if not thinking: thinking = True print("=============Start Reasoning=============") print(getattr(choice.delta, "reasoning_content"), end="") if choice.delta and choice.delta.content: if thinking: thinking = False print("\n=============End Reasoning=============") print(choice.delta.content, end="") ``` ### Call kimi-k2.6: reasoning output by default `kimi-k2.6` is the general-purpose thinking model. Thinking is enabled by default, so the basic call below outputs reasoning content without passing the `thinking` parameter (to disable thinking or enable Preserved Thinking, see [the thinking parameter](#thinking-parameter) below): ```bash theme={null} $ curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k2.6", "messages": [ { "role": "system", "content": "You are Kimi." }, { "role": "user", "content": "Please explain why 1+1=2." } ] }' ``` ```python theme={null} import os import openai client = openai.Client( base_url="https://api.moonshot.ai/v1", api_key=os.getenv("MOONSHOT_API_KEY"), ) stream = client.chat.completions.create( model="kimi-k2.6", messages=[ { "role": "system", "content": "You are Kimi.", }, { "role": "user", "content": "Please explain why 1+1=2." }, ], max_tokens=1024*32, stream=True, # temperature is not modifiable, so no need to set it; thinking is enabled by default, no extra parameters needed ) thinking = False for chunk in stream: if chunk.choices: choice = chunk.choices[0] if choice.delta and hasattr(choice.delta, "reasoning_content"): if not thinking: thinking = True print("=============Start Reasoning=============") print(getattr(choice.delta, "reasoning_content"), end="") if choice.delta and choice.delta.content: if thinking: thinking = False print("\n=============End Reasoning=============") print(choice.delta.content, end="") ``` ## Control thinking behavior ### K3: adjust reasoning effort with `reasoning_effort` `kimi-k3` always reasons and does not support the `thinking` parameter. Adjust reasoning effort with the top-level `reasoning_effort` field (`"low"` / `"high"` / `"max"`, default `"max"`); see [Reasoning Effort](/docs/guide/use-reasoning-effort) for usage and examples. ### Control kimi-k2.6 thinking with the thinking parameter `kimi-k2.6` controls thinking behavior via the `thinking` parameter, which has two sub-fields: * `thinking.type`: `"enabled"` (default) | `"disabled"` — controls whether thinking is on. Since it defaults to `"enabled"`, the example above thinks without passing it explicitly; for a disable example see [Disable Thinking Capability Example](/docs/guide/kimi-k2-6-quickstart#disable-thinking-capability-example). * `thinking.keep`: `null` (default, ignores historical turns' thinking) | `"all"` (keeps previous turns' `reasoning_content`, enabling Preserved Thinking — see [Preserved Thinking](#preserved-thinking) for usage). ## Read reasoning\_content from the response For thinking models such as `kimi-k2.7-code` and `kimi-k2.6` (with thinking enabled), the API response carries the model's reasoning in the `reasoning_content` field. When reading this field: * In the OpenAI SDK, `ChoiceDelta` and `ChatCompletionMessage` types do not provide a `reasoning_content` field directly, so you cannot access it via `.reasoning_content`. You must use `hasattr(obj, "reasoning_content")` to check if the field exists, and if so, use `getattr(obj, "reasoning_content")` to retrieve its value. * If you use other frameworks or directly interface with the HTTP API, you can directly obtain the `reasoning_content` field at the same level as the `content` field. * In streaming output (`stream=True`), the `reasoning_content` field will always appear before the `content` field. In your business logic, you can detect if the `content` field has been output to determine if the reasoning (inference process) is finished. * Tokens in `reasoning_content` are also controlled by the `max_tokens` parameter: the sum of tokens in `reasoning_content` and `content` must be less than or equal to `max_tokens`. ## Configure multi-step tool calls `kimi-k2.7-code` and `kimi-k2.6` (with thinking enabled) are designed to perform deep reasoning across multiple tool calls, enabling them to tackle highly complex tasks. To get reliable results, **always follow these configuration rules when using thinking models:** * Within a single task (the multi-step reasoning produced during one tool-call loop), keep all of the reasoning content from the context (the `reasoning_content` field) and send it back with the request; the model will choose which parts are necessary and forward them for reasoning. Whether historical thinking is preserved *across turns* is controlled by `thinking.keep` (`kimi-k2.6` defaults to `null` and does not keep it, while `kimi-k2.7-code` always keeps it). * Set `max_tokens >= 16000` to ensure the full `reasoning_content` and `content` can be returned without truncation. * **Do not set `temperature`.** For `kimi-k2.7-code` and `kimi-k2.6`, `temperature` is not modifiable — use the default and do not pass it explicitly (see [Model Parameter Reference](/docs/api/models-overview)). * Enable streaming (`stream=True`). Because thinking models return both `reasoning_content` and regular `content`, the response is larger than usual. Streaming delivers a better user experience and helps avoid network-timeout issues. ### Complete example: generate a daily news report The example below demonstrates a "Daily News Report Generation" scenario. The model will sequentially call official tools like `date` (to get the date) and `web_search` (to search today's news), and will present deep reasoning throughout this process: ```python expandable theme={null} import os import json import httpx import openai class FormulaChatClient: def __init__(self, base_url: str, api_key: str): """Initialize Formula client""" self.base_url = base_url self.api_key = api_key self.openai = openai.Client( base_url=base_url, api_key=api_key, ) self.httpx = httpx.Client( base_url=base_url, headers={"Authorization": f"Bearer {api_key}"}, timeout=30.0, ) # Using kimi-k2.6 model. Thinking is enabled by default self.model = "kimi-k2.6" def get_tools(self, formula_uri: str): """Get tool definitions from Formula API""" response = self.httpx.get(f"/formulas/{formula_uri}/tools") response.raise_for_status() try: return response.json().get("tools", []) except json.JSONDecodeError as e: print(f"Error: Unable to parse JSON (status code: {response.status_code})") print(f"Response content: {response.text[:500]}") raise def call_tool(self, formula_uri: str, function: str, args: dict): """Call an official tool""" response = self.httpx.post( f"/formulas/{formula_uri}/fibers", json={"name": function, "arguments": json.dumps(args)}, ) response.raise_for_status() fiber = response.json() if fiber.get("status", "") == "succeeded": return fiber["context"].get("output") or fiber["context"].get("encrypted_output") if "error" in fiber: return f"Error: {fiber['error']}" if "error" in fiber.get("context", {}): return f"Error: {fiber['context']['error']}" return "Error: Unknown error" def close(self): """Close the client connection""" self.httpx.close() # Initialize client base_url = os.getenv("MOONSHOT_BASE_URL", "https://api.moonshot.ai/v1") api_key = os.getenv("MOONSHOT_API_KEY") if not api_key: raise ValueError("MOONSHOT_API_KEY environment variable not set. Please set your API key.") print(f"Base URL: {base_url}") print(f"API Key: {api_key[:10]}...{api_key[-10:] if len(api_key) > 20 else api_key}\n") client = FormulaChatClient(base_url, api_key) # Define the official tool Formula URIs to use formula_uris = [ "moonshot/date:latest", "moonshot/web-search:latest" ] # Load all tool definitions and build mapping print("Loading official tools...") all_tools = [] tool_to_uri = {} # function.name -> formula_uri for uri in formula_uris: try: tools = client.get_tools(uri) for tool in tools: func = tool.get("function") if func: func_name = func.get("name") if func_name: tool_to_uri[func_name] = uri all_tools.append(tool) print(f" Loaded tool: {func_name} from {uri}") except Exception as e: print(f" Warning: Failed to load tool {uri}: {e}") continue print(f"Loaded {len(all_tools)} tools in total\n") if not all_tools: raise ValueError("No tools loaded. Please check API key and network connection.") # Initialize message list messages = [ { "role": "system", "content": "You are Kimi, a professional news analyst. You excel at collecting, analyzing, and organizing information to generate high-quality news reports.", }, ] # User request to generate today's news report user_request = "Please help me generate a daily news report including important technology, economy, and society news." messages.append({ "role": "user", "content": user_request }) print(f"User request: {user_request}\n") # Begin multi-step conversation loop max_iterations = 10 # Prevent infinite loops for iteration in range(max_iterations): try: completion = client.openai.chat.completions.create( model=client.model, messages=messages, max_tokens=1024 * 32, tools=all_tools, ) except openai.AuthenticationError as e: print(f"Authentication error: {e}") print("Please check if the API key is correct and has the required permissions") raise except Exception as e: print(f"Error while calling the model: {e}") raise # Get response message = completion.choices[0].message # Print reasoning process if hasattr(message, "reasoning_content"): print(f"=============Reasoning round {iteration + 1} starts=============") reasoning = getattr(message, "reasoning_content") if reasoning: print(reasoning[:500] + "..." if len(reasoning) > 500 else reasoning) print(f"=============Reasoning round {iteration + 1} ends=============\n") # Add assistant message to context (preserve reasoning_content) messages.append(message) # If the model did not call any tools, conversation is done if not message.tool_calls: print("=============Final Answer=============") print(message.content) break # Handle tool calls print(f"The model decided to call {len(message.tool_calls)} tool(s):\n") for tool_call in message.tool_calls: func_name = tool_call.function.name args = json.loads(tool_call.function.arguments) print(f"Calling tool: {func_name}") print(f"Arguments: {json.dumps(args, ensure_ascii=False, indent=2)}") # Get corresponding formula_uri formula_uri = tool_to_uri.get(func_name) if not formula_uri: print(f"Error: Could not find Formula URI for tool {func_name}") continue # Call the tool result = client.call_tool(formula_uri, func_name, args) # Print result (truncate if too long) if len(str(result)) > 200: print(f"Tool result: {str(result)[:200]}...\n") else: print(f"Tool result: {result}\n") # Add tool result to message list tool_message = { "role": "tool", "tool_call_id": tool_call.id, "name": func_name, "content": result } messages.append(tool_message) print("\nConversation completed!") # Cleanup client.close() ``` This process demonstrates how thinking models such as `kimi-k2.7-code` and `kimi-k2.6` (with thinking enabled) use deep reasoning to plan and execute complex multi-step tasks, with detailed reasoning steps (`reasoning_content`) preserved in the context to ensure accurate tool use at every stage. ## Preserve thinking across turns (Preserved Thinking) Preserved Thinking means passing the `reasoning_content` of previous turns through to the model in a multi-turn conversation, so that the model can continue its prior chain of thought when reasoning in the current turn. For `kimi-k2.6`, use the `thinking.keep` parameter in the request body to control whether historical thinking is preserved: | Value | Behavior | | -------------------------- | ------------------------------------------------------------------------------- | | `null` / omitted (default) | Historical `reasoning_content` is ignored. Shorter context and lower cost. | | `"all"` | Historical `reasoning_content` is fully preserved, enabling Preserved Thinking. | `thinking.keep` only affects `reasoning_content` from historical turns; it does **not** change whether the model generates/outputs thinking content within the current turn (that is controlled by `thinking.type`). Recommended to use `keep: "all"` together with `type: "enabled"`. For `kimi-k2.7-code`, Preserved Thinking is always on and cannot be turned off: `thinking.keep` is treated as `"all"` whether you omit it or pass the only valid value `"all"` (passing any other invalid value returns an error). When using this model you must therefore (not optionally) keep the `reasoning_content` of historical assistant messages in `messages` as-is, exactly as shown in the example below. When using `keep: "all"`, keep the `reasoning_content` from every historical assistant message in `messages` as-is. The simplest way is to append the assistant message returned from the previous API call directly back into `messages`, as shown below: ```bash theme={null} $ curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k2.6", "messages": [ {"role": "system", "content": "You are Kimi."}, {"role": "user", "content": "First question..."}, { "role": "assistant", "reasoning_content": "", "content": "" }, {"role": "user", "content": "Please continue the analysis and derive the next step."} ], "thinking": { "type": "enabled", "keep": "all" } }' ``` ```python theme={null} import os import openai client = openai.Client( base_url="https://api.moonshot.ai/v1", api_key=os.getenv("MOONSHOT_API_KEY"), ) # Keep the assistant message (including reasoning_content) from every previous API call in messages messages = [ {"role": "system", "content": "You are Kimi."}, {"role": "user", "content": "First question..."}, { "role": "assistant", "reasoning_content": "", "content": "", }, {"role": "user", "content": "Please continue the analysis and derive the next step."}, ] response = client.chat.completions.create( model="kimi-k2.6", messages=messages, stream=True, extra_body={"thinking": {"type": "enabled", "keep": "all"}}, ) ``` `reasoning_content` counts toward token consumption. When Preserved Thinking is enabled, historical thinking content keeps occupying the context window and is billed accordingly. Use it wisely. ## Frequently Asked Questions ### Q1: Why should I keep `reasoning_content`? A: Keeping `reasoning_content` ensures continuity in multi-step reasoning, especially during tool calls. Pass the complete assistant message returned by the API back in `messages` as-is. For K3, this is required in multi-turn conversations and tool-call loops. For K2.x, cross-turn preservation follows each model's `thinking.keep` behavior: `kimi-k2.6` does not preserve it by default, while `kimi-k2.7-code` always does. ### Q2: Does `reasoning_content` consume extra tokens? A: Yes, `reasoning_content` counts towards your input/output token quota. For detailed pricing, see [Pricing](/docs/pricing/chat). # Tool Choice Source: https://platform.kimi.ai/docs/guide/use-tool-choice Use `tool_choice` to let Kimi select tools automatically, require a tool call, forbid tools, or force a specific function. Once tools are declared (via `tools`), the model decides on its own whether the current turn needs a tool call. The `tool_choice` parameter gives you explicit control over this behavior: force a call, forbid calls entirely, or keep the default. ## Force a tool call: `"required"` Use this when your workflow must go through the tool path — for example, mandatory retrieval or a mandatory database lookup — and the model is not allowed to answer from memory: ```json theme={null} { "tool_choice": "required" } ``` The model must call at least one tool in this turn. Make sure the request declares at least one callable tool. A typical use is the tool-search pattern: set `"required"` on the first turn to force the model to call `search_tools`, then switch back to `"auto"` after retrieval — see [Kimi K3 API Tool Calling Best Practices](/docs/guide/kimi-k3-tool-calling-best-practice). ## Forbid tool calls: `"none"` Use this when the request only needs a plain-text answer and you don't want the model to trigger a tool call by mistake: ```json theme={null} { "tool_choice": "none" } ``` The model replies with plain text and produces no `tool_calls`, which also reduces latency and token consumption. ## Let the model decide: `"auto"` (default) Omitting `tool_choice` is the same as `"auto"`: the model decides based on the context whether to call a tool. This fits regular conversations. ## Force a specific tool: pass a function object Beyond the three enum values, `tool_choice` also accepts a function object that forces the model to call the specified tool: ```json theme={null} { "tool_choice": {"type": "function", "function": {"name": "get_weather"}} } ``` Forcing a specific tool is currently incompatible with thinking: with thinking enabled, the request returns a 400 error (`tool_choice 'specified' is incompatible with thinking enabled`). ## Full request example The following example declares a weather tool and uses `tool_choice: "required"` to force the model to call it: ```bash theme={null} $ curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k3", "messages": [ { "role": "user", "content": "What is the weather like in Beijing today?" } ], "tools": [ { "type": "function", "function": { "name": "get_weather", "description": "Get the current weather for a given city", "parameters": { "type": "object", "properties": { "city": { "type": "string", "description": "City name" } }, "required": ["city"] } } } ], "tool_choice": "required" }' ``` ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "user", "content": "What is the weather like in Beijing today?"}, ], tools=[ { "type": "function", "function": { "name": "get_weather", "description": "Get the current weather for a given city", "parameters": { "type": "object", "properties": { "city": {"type": "string", "description": "City name"} }, "required": ["city"], }, }, } ], # Force the model to call at least one tool; defaults to "auto" when omitted tool_choice="required", ) print(completion.choices[0].message.tool_calls) ``` ## Notes * `tool_choice` is a request-level parameter: it takes effect independently for each request and only constrains tool selection for that generation; * Setting `tool_choice` (or not setting it) **does not invalidate the prefix cache**, so feel free to adjust it on a per-request basis. ## Related reading * [Kimi K3 API Tool Calling Best Practices](/docs/guide/kimi-k3-tool-calling-best-practice): combined practices for dynamic loading, tool\_choice, and reasoning effort * [Dynamically Loaded Tools](/docs/guide/use-dynamic-tool-loading): inject tool definitions on demand when you have a large tool inventory, reducing token usage and improving tool-selection accuracy * [Use Kimi API for Tool Calls](/docs/guide/use-kimi-api-to-complete-tool-calls): the complete tool-calling workflow and examples * [Model Parameter Reference](/docs/api/models-overview): per-model support for the `tool_choice` parameter # Use Kimi API's Internet Search Functionality Source: https://platform.kimi.ai/docs/guide/use-web-search Add web search to Kimi API applications using the recommended official tool channel or the built-in `$web_search` flow for supported models. When using web search with `kimi-k3`, we recommend the Formula API official tools channel (OpenAI protocol, standard `function` tool); see [How to Use Official Tools in Kimi API](/docs/guide/use-official-tools). `$web_search` (of type `builtin_function`) is Kimi's built-in web search tool function, implemented on top of the `tool_calls` usage: the model only generates the search arguments, while the search itself is defined and executed by the Kimi large language model. When you don't want to implement search-engine calls, page fetching, and content cleanup yourself, declare this built-in tool to get out-of-the-box web search. Its basic usage and flow are the same as a regular `tool_calls` tool call — define the tool, submit it via `tools`, let the model generate the arguments, return the execution result, and get the model's reply. For the full flow, see [Use Kimi API to Complete Tool Calls](/docs/guide/use-kimi-api-to-complete-tool-calls); this page only highlights where `$web_search` differs from a regular `function`. ## Declare `$web_search` Unlike an ordinary `tool`, the `$web_search` function does not require specific parameter descriptions — declaring only `type` and `function.name` in `tools` is enough to register it: ```python theme={null} tools = [ { "type": "builtin_function", # <-- We use builtin_function to indicate Kimi built-in tools, which also distinguishes it from ordinary function "function": { "name": "$web_search", }, }, ] ``` **The `$web_search` function is prefixed with a dollar sign `$`, which is our agreed way to indicate Kimi built-in functions** (in ordinary `function` definitions, the dollar sign `$` is not allowed), and if there are other Kimi built-in functions in the future, they will also be prefixed with the dollar sign `$`. **`$web_search` works directly with each model's reasoning behavior**: `kimi-k3` always reasons, and `kimi-k2.6` can also perform web search with thinking enabled. `$web_search` can coexist with other ordinary `function` tools: within the same `tools` declaration, you can freely mix tools with `type=builtin_function` and `type=function`. ## Run the web search When using the `$web_search` function, the basic flow is no different from that of a regular `function` — developers don't even need to modify the original code for executing `tool_calls`. The following example shows the complete flow: declare `$web_search`, ask a question, and loop over `tool_calls` until the model returns its final reply, where `search_impl` simply returns the model-generated arguments as-is: The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ```python theme={null} from typing import * import os import json from openai import OpenAI from openai.types.chat.chat_completion import Choice client = OpenAI( base_url="https://api.moonshot.ai/v1", api_key=os.environ.get("MOONSHOT_API_KEY"), ) # Specific implementation of the search tool, here we only need to return the arguments def search_impl(arguments: Dict[str, Any]) -> Any: """ When using the search tool provided by Moonshot AI, you only need to return the arguments as-is, no additional processing logic is needed. But if you want to use other models and retain the web search functionality, you only need to modify the implementation here (e.g., calling search and getting web content, etc.), the function signature remains unchanged and still works. This maximizes compatibility, allowing you to switch between different models without needing destructive changes to the code. """ return arguments def chat(messages) -> Choice: completion = client.chat.completions.create( model="kimi-k3", messages=messages, max_tokens=32768, tools=[ { "type": "builtin_function", # <-- Use builtin_function to declare the $web_search function, please include the complete tools declaration in every request "function": { "name": "$web_search", }, } ] ) return completion.choices[0] def main(): messages = [ {"role": "system", "content": "You are Kimi."}, ] # Initial question messages.append({ "role": "user", "content": "Please search for Moonshot AI Context Caching technology and tell me what it is." }) finish_reason = None while finish_reason is None or finish_reason == "tool_calls": choice = chat(messages) finish_reason = choice.finish_reason if finish_reason == "tool_calls": # <-- Determine whether the current returned content contains tool_calls messages.append(choice.message) # <-- We also add the assistant message returned by the Kimi model to the context, so that the Kimi model can understand our request in the next request for tool_call in choice.message.tool_calls: # <-- tool_calls may be multiple, so we use a loop to execute them one by one tool_call_name = tool_call.function.name tool_call_arguments = json.loads(tool_call.function.arguments) # <-- arguments is a serialized JSON Object, we need to use json.loads to deserialize it if tool_call_name == "$web_search": tool_result = search_impl(tool_call_arguments) else: tool_result = f"Error: unable to find tool by name '{tool_call_name}'" # Construct a role=tool message using the function execution result to show the model the tool call result; # Note that we need to provide tool_call_id and name fields in the message so that the Kimi model # can correctly match the corresponding tool_call. messages.append({ "role": "tool", "tool_call_id": tool_call.id, "name": tool_call_name, "content": json.dumps(tool_result), # <-- We agree to submit tool call results to the Kimi model in string format, so here we use json.dumps to serialize the execution result into a string }) print(choice.message.content) # <-- Here, we return the reply generated by the model to the user if __name__ == '__main__': main() ``` ```js theme={null} const openai = require('openai'); // Need to install openai library const client = new openai.OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }); const tools = [ { "type": "builtin_function", "function": { "name": "$web_search", }, } ]; function search_impl(args) { return args } const messages = [ { "role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are better at Chinese and English conversations. You will provide users with safe, helpful, and accurate answers. At the same time, you will refuse to answer any questions involving terrorism, racial discrimination, pornography, and violence. Moonshot AI is a proper noun and cannot be translated into other languages." }, { "role": "user", "content": "Please search for the China A-share index on October 8, 2024?" } // Ask Kimi model to search online in the question ]; let finishReason = null; async function main() { while (finishReason === null || finishReason === "tool_calls") { const completion = await client.chat.completions.create({ model: "kimi-k3", messages: messages, tools: tools // <-- We submit the defined tools to the Kimi model through the tools parameter }); const choice = completion.choices[0]; console.log(choice); finishReason = choice.finish_reason; console.log(finishReason); if (finishReason === "tool_calls") { // <-- Check if the returned content contains tool_calls messages.push(choice.message); // <-- We add the assistant message returned by the Kimi model to the context so that the Kimi model can understand our request in the next request for (const toolCall of choice.message.tool_calls) { // <-- tool_calls may be multiple, so we use a loop to execute them one by one const tool_call_name = toolCall.function.name; const tool_call_arguments = JSON.parse(toolCall.function.arguments); // <-- arguments is a serialized JSON Object, we need to use JSON.parse to deserialize it let tool_result; if (tool_call_name == "$web_search") { tool_result = search_impl(tool_call_arguments) } else { tool_result = 'no tool found' } // Construct a message with role=tool using the function execution result to show the model the tool call result; // Note that we need to provide tool_call_id and name fields in the message so that the Kimi model // can correctly match the corresponding tool_call. console.log("toolCall.id"); console.log(toolCall.id); console.log("tool_call_name"); console.log(tool_call_name); console.log("tool_result"); console.log(tool_result); messages.push({ "role": "tool", "tool_call_id": toolCall.id, "name": tool_call_name, "content": JSON.stringify(tool_result), // <-- We agree to submit tool call results to the Kimi model in string format, so we use JSON.stringify to serialize the execution result into a string }); } } console.log(choice.message.content); // <-- Here, we return the model-generated reply to the user } } main(); ``` Why doesn't `search_impl` need any logic for searching, parsing, or obtaining web content? As the name `builtin_function` suggests, `$web_search` is a built-in function of the Kimi large language model: it is defined by the Kimi large language model and executed by it as well. 1. When Kimi large language model generates a response with `finish_reason=tool_calls`, it means that Kimi large language model has realized that it needs to execute the `$web_search` function and has already prepared everything for it; 2. Kimi large language model will return the necessary parameters for executing the function in the form of `tool_call.function.arguments`. However, these parameters are not executed by the caller. The caller just needs to submit `tool_call.function.arguments` to Kimi large language model as they are, and Kimi large language model will execute the corresponding online search process; 3. When the user submits `tool_call.function.arguments` using a `message` with `role=tool`, Kimi large language model will immediately start the online search process and generate a readable message for the user based on the search and reading results, which is a `message` with `finish_reason=stop`; ## Switch to your own search implementation The online search function provided by the Kimi API aims to offer a reliable large language model online search solution without breaking the compatibility of the original API and SDK, and it is fully compatible with the original `tool_calls` feature of the Kimi large language model. **If you want to switch from Kimi's online search function to your own implementation, you can do so in just two simple steps without disrupting the overall structure of your code:** 1. Modify the `tool` definition of `$web_search` to your own implementation (including `name`, `description` etc.). You may need to add additional information in `tool.function` to inform the model of the specific parameters it needs to generate. You can add any parameters you need in the `parameters` field; 2. Change the implementation of the `search_impl` function: when using Kimi's `$web_search`, you just need to return the input `arguments` as they are; if you use your own online search service, you may need to fully implement the `search` and `crawl` functions — call search engine APIs (or implement your own content search) to retrieve URLs and summaries, fetch web page content based on URLs (which might require different reading rules for different websites), clean and organize the fetched web page content into a format that the model can easily recognize, such as Markdown, and handle various errors and exceptions, such as no search results or failure to fetch web page content; After completing the above steps, you will have successfully migrated from Kimi's online search function to your own implementation. ## Track web search token usage When using the `$web_search` function provided by Kimi, the search results are also counted towards the tokens occupied by the prompt (i.e., `prompt_tokens`). Typically, since the results of web searches contain a lot of content, the token consumption can be quite high. To avoid unknowingly using up a large number of tokens, an extra `total_tokens` field is added under the `usage` object inside the generated `arguments` (read it as `arguments.usage.total_tokens`), informing the caller of the total number of tokens occupied by the search content. These tokens will be included in the `prompt_tokens` once the entire web search process is completed. The following example shows how to read the `total_tokens` occupied by the search results, along with the token consumption of the whole conversation: ```python theme={null} from typing import * import os import json from openai import OpenAI from openai.types.chat.chat_completion import Choice client = OpenAI( base_url="https://api.moonshot.ai/v1", api_key=os.environ.get("MOONSHOT_API_KEY"), ) # Specific implementation of search tool, here we only need to return parameters def search_impl(arguments: Dict[str, Any]) -> Any: """ When using search tool provided by Moonshot AI, you only need to return arguments as is, without additional processing logic. But if you want to use other models and retain web search function, you only need to modify implementation here (such as calling search and getting web content, etc.), function signature remains same, still work. This maximizes compatibility, allowing you to switch between different models without breaking changes to code. """ return arguments def chat(messages) -> Choice: completion = client.chat.completions.create( model="kimi-k3", messages=messages, max_tokens=32768, tools=[ { "type": "builtin_function", "function": { "name": "$web_search", }, } ] ) usage = completion.usage choice = completion.choices[0] # ========================================================================= # By judging finish_reason = stop, we print out the tokens consumed after completing the web search process if choice.finish_reason == "stop": print(f"chat_prompt_tokens: {usage.prompt_tokens}") print(f"chat_completion_tokens: {usage.completion_tokens}") print(f"chat_total_tokens: {usage.total_tokens}") # ========================================================================= return choice def main(): messages = [ {"role": "system", "content": "You are Kimi."}, ] # Initial question messages.append({ "role": "user", "content": "Please search for Moonshot AI Context Caching technology and tell me what it is." }) finish_reason = None while finish_reason is None or finish_reason == "tool_calls": choice = chat(messages) finish_reason = choice.finish_reason if finish_reason == "tool_calls": # Append the complete assistant message, preserving reasoning_content and tool_calls unchanged. messages.append(choice.message) for tool_call in choice.message.tool_calls: tool_call_name = tool_call.function.name tool_call_arguments = json.loads( tool_call.function.arguments) if tool_call_name == "$web_search": # =================================================================== # We print out the tokens generated by web search results during the web search process search_content_total_tokens = tool_call_arguments.get("usage", {}).get("total_tokens") print(f"search_content_total_tokens: {search_content_total_tokens}") # =================================================================== tool_result = search_impl(tool_call_arguments) else: tool_result = f"Error: unable to find tool by name '{tool_call_name}'" messages.append({ "role": "tool", "tool_call_id": tool_call.id, "name": tool_call_name, "content": json.dumps(tool_result), }) print(choice.message.content) if __name__ == '__main__': main() ``` Running the above code yields the following output: ```shell theme={null} search_content_total_tokens: 13046 # <-- This represents the number of tokens occupied by the web search results due to the web search action. chat_prompt_tokens: 13212 # <-- This represents the number of input tokens, including the web search results. chat_completion_tokens: 295 # <-- This represents the number of tokens generated by the Kimi large language model based on the web search results. chat_total_tokens: 13507 # <-- This represents the total number of tokens consumed, including the web search process. # The content generated by the Kimi large language model based on the web search results is omitted here. ``` ## About Model Size Selection Enabling web search significantly increases context length because search results are appended to the conversation. To avoid triggering `Input token length too long`, we recommend using `kimi-k3`, which has a 1M-token context window: ```python theme={null} def chat(messages) -> Choice: completion = client.chat.completions.create( model="kimi-k3", messages=messages, tools=[ { "type": "builtin_function", # <-- Use builtin_function to declare the $web_search function. Please include the full tools declaration in each request. "function": { "name": "$web_search", }, } ] ) return completion.choices[0] ``` ## Web search billing In addition to token consumption, we also charge a call fee for each web search. For details, see [Pricing](/docs/pricing/tools). # Use the Streaming Feature of the Kimi API Source: https://platform.kimi.ai/docs/guide/utilize-the-streaming-output-feature-of-kimi-api Use Kimi API server-sent events to reduce time to first token, parse incremental output, and capture final usage data. After receiving a question, the Kimi large language model first performs inference and then generates the answer one Token at a time. Streaming sends Tokens to the client as soon as a certain number of them (usually 1 Token) is generated, instead of waiting until the full response is complete. Waiting for the complete response usually takes several seconds — for complex questions and long replies it can stretch to 10 or even 20 seconds; with streaming, users see the first Token immediately, which significantly reduces wait time. When you chat with [Kimi AI Assistant](https://kimi.ai), the reply appears character by character — that is streaming in action. ## Enable Streaming Output Set `stream=True` in the request to enable streaming. The SDK then returns an iterable — loop over it to read data chunks one by one. Each chunk has a structure similar to a completion, except the `message` field is replaced by a `delta` field: The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key = os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url = "https://api.moonshot.ai/v1", ) stream = client.chat.completions.create( model = "kimi-k3", messages = [ {"role": "system", "content": "You are Kimi, an artificial intelligence assistant provided by Moonshot AI, who is better at conversing in Chinese and English. You provide users with safe, helpful, and accurate answers. At the same time, you refuse to answer any questions related to terrorism, racism, pornography, and violence. Moonshot AI is a proper noun and should not be translated into other languages."}, {"role": "user", "content": "Hello, my name is Li Lei, what is 1+1?"} ], stream=True, # <-- Note here, we enable streaming output mode by setting stream=True ) # When streaming output mode is enabled (stream=True), the content returned by the SDK also changes. We no longer directly access the choice in the return value # Instead, we access each individual chunk in the return value through a for loop for chunk in stream: # Here, the structure of each chunk is similar to the previous completion, but the message field is replaced with the delta field delta = chunk.choices[0].delta # <-- The message field is replaced with the delta field if delta.content: # When printing the content, since it is streaming output, to ensure the coherence of the sentence, we do not add # line breaks manually, so we set end="" to cancel the line break of print. print(delta.content, end="") ``` ```js theme={null} const OpenAI = require('openai') const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, // Set the MOONSHOT_API_KEY environment variable before running this example baseURL: "https://api.moonshot.ai/v1", }) async function main() { const stream = await client.chat.completions.create({ model: "kimi-k3", messages: [ {role: "system", content: "You are Kimi, an artificial intelligence assistant provided by Moonshot AI, who is better at conversing in Chinese and English. You provide users with safe, helpful, and accurate answers. At the same time, you refuse to answer any questions related to terrorism, racism, pornography, and violence. Moonshot AI is a proper noun and should not be translated into other languages."}, {role: "user", content: "Hello, my name is Li Lei, what is 1+1?"} ], stream: true, // <-- Note here, we enable streaming output mode by setting stream=True }) // When streaming output mode is enabled (stream=True), the content returned by the SDK also changes. We no longer directly access the choice in the return value // Instead, we access each individual chunk in the return value through a for loop for await (chunk of stream) { // Here, the structure of each chunk is similar to the previous completion, but the message field is replaced with the delta field delta = chunk.choices[0].delta // <-- The message field is replaced with the delta field if (delta.content) { // When printing the content, since it is streaming output, to ensure the coherence of the sentence, we do not add // line breaks manually, so we set end="" to cancel the line break of print. console.log(delta.content, end="") } } } main() ``` ## Parse the SSE Response Body With streaming enabled, the API no longer returns a JSON response (`Content-Type: application/json`); it returns `Content-Type: text/event-stream` (SSE) instead, which lets the server continuously push Tokens to the client. An [SSE](https://kimi.ai/share/cr7boh3dqn37a5q9tds0) response body looks like this: ```text theme={null} data: {"id":"cmpl-1305b94c570f447fbde3180560736287","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k3","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]} data: {"id":"cmpl-1305b94c570f447fbde3180560736287","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k3","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]} ... data: {"id":"cmpl-1305b94c570f447fbde3180560736287","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k3","choices":[{"index":0,"delta":{"content":"."},"finish_reason":null}]} data: {"id":"cmpl-1305b94c570f447fbde3180560736287","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k3","choices":[{"index":0,"delta":{},"finish_reason":"stop","usage":{"prompt_tokens":19,"completion_tokens":13,"total_tokens":32}}]} data: [DONE] ``` In the response body, each data chunk starts with the `data: ` prefix, followed by a valid JSON object, and ends with two newline characters `\n\n`. Once all chunks are transmitted, the server sends `data: [DONE]` to mark completion, at which point you can close the connection. *Note: always use `data: [DONE]` to determine whether the data has been fully transmitted, not `finish_reason` or any other means. If you have not received `data: [DONE]`, do not consider the transmission complete even if `finish_reason=stop` was received; in other words, until `data: [DONE]` arrives, the message should be considered **incomplete**.* During streaming, the `content` field is delivered chunk by chunk; `role` and `usage` are not repeated in every chunk — `role` appears only in the first chunk, and `usage` only in the last one. ## Count Token Usage There are two ways to count tokens. The most direct and accurate one is to wait until all chunks have been transmitted, then read the `usage` field of the last chunk to see the request's `prompt_tokens`/`completion_tokens`/`total_tokens`: ```text theme={null} ... data: {"id":"cmpl-1305b94c570f447fbde3180560736287","object":"chat.completion.chunk","created":1698999575,"model":"kimi-k3","choices":[{"index":0,"delta":{},"finish_reason":"stop","usage":{"prompt_tokens":19,"completion_tokens":13,"total_tokens":32}}]} ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ Check the number of tokens generated by the current request through the usage field in the last data chunk data: [DONE] ``` Note that `usage` is nested inside `choices[0]` of the last chunk (`choices[0].usage`), not at the top level of the chunk. With the OpenAI SDK, `chunk.usage` is `None` — read `chunk.choices[0].usage` instead, or parse the raw SSE chunks directly. However, a stream can be interrupted by uncontrollable factors such as a network drop or a client-side error, in which case the last chunk never arrives and the request's token consumption cannot be determined. To avoid this, save the content of every chunk you receive and, once the request ends (whether successfully or not), call the token-count endpoint to compute the actual consumption: ```python theme={null} import os import httpx from openai import OpenAI client = OpenAI( api_key = os.environ["MOONSHOT_API_KEY"], # Set the MOONSHOT_API_KEY environment variable before running this example base_url = "https://api.moonshot.ai/v1", ) stream = client.chat.completions.create( model = "kimi-k3", messages = [ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI, who excels in Chinese and English conversations. You provide users with safe, helpful, and accurate answers while rejecting any questions related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated."}, {"role": "user", "content": "Hello, my name is Li Lei. What is 1+1?"} ], stream=True, # <-- Note here, we enable streaming output mode by setting stream=True ) def estimate_token_count(input: str) -> int: """ Implement your token calculation logic here, or directly call our token calculation interface to compute tokens. https://api.moonshot.ai/v1/tokenizers/estimate-token-count """ header = { "Authorization": f"Bearer {os.environ['MOONSHOT_API_KEY']}", } data = { "model": "kimi-k3", "messages": [ {"role": "user", "content": input}, ] } r = httpx.post("https://api.moonshot.ai/v1/tokenizers/estimate-token-count", headers=header, json=data) r.raise_for_status() return r.json()["data"]["total_tokens"] completion = [] for chunk in stream: delta = chunk.choices[0].delta if delta.content: completion.append(delta.content) print("completion_tokens:", estimate_token_count("".join(completion))) ``` ```js theme={null} const axios = require('axios'); const OpenAI = require('openai'); client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }) async function estimate_token_count(input_messages) { /* Implement your token calculation logic here, or directly call our token calculation interface to compute tokens. https://api.moonshot.ai/v1/tokenizers/estimate-token-count */ header = { "Authorization": `Bearer ${process.env.MOONSHOT_API_KEY}`, } data = { "model": "kimi-k3", "messages": input_messages, } r = await axios.post("https://api.moonshot.ai/v1/tokenizers/estimate-token-count", data, {headers: header}) .catch(function (error) { console.log(error) }) return r.data.data.total_tokens } async function main() { const stream = await client.chat.completions.create({ model: "kimi-k3", messages: [ {role: "system", content: "You are Kimi, an AI assistant provided by Moonshot AI, who excels in Chinese and English conversations. You provide users with safe, helpful, and accurate answers while rejecting any questions related to terrorism, racism, or explicit content. Moonshot AI is a proper noun and should not be translated."}, {role: "user", content: "Hello, my name is Li Lei. What is 1+1?"} ], stream: true, // <-- Note here, we enable streaming output mode by setting stream=True }) const completion = []; for await (chunk of stream) { const delta = chunk.choices[0].delta if (delta.content) { completion.push(delta.content) } } console.log("completion_tokens:", await estimate_token_count(completion.join(""))) } main() ``` ## Stop Streaming Output To terminate the output early, simply close the HTTP connection or discard subsequent chunks — for example, `break` out of the loop: ```python theme={null} for chunk in stream: if condition: break ``` ## Handle SSE Without an SDK In a language without an SDK, or when the SDK cannot accommodate your business logic, you can interface with the HTTP API directly to handle streaming output. The following examples show how to read and parse the [SSE](https://kimi.ai/share/cr7boh3dqn37a5q9tds0) response body line by line; see the code comments for details: ```python theme={null} import os import json import httpx # We use the httpx library to make our HTTP requests data = { "model": "kimi-k3", "messages": [ # Specific messages ], "stream": True, } # Use httpx to send a chat request to the Kimi large language model and get the response r r = httpx.post("https://api.moonshot.ai/v1/chat/completions", headers={"Authorization": f"Bearer {os.environ['MOONSHOT_API_KEY']}"}, json=data) if r.status_code != 200: raise Exception(r.text) data: str # Here, we use the iter_lines method to read the response body line by line for line in r.iter_lines(): # Remove leading and trailing spaces from each line to better handle data chunks line = line.strip() # Next, we need to handle three different cases: # 1. If the current line is empty, it indicates that the previous data chunk has been received (as mentioned earlier, the data chunk transmission ends with two newline characters), we can deserialize the data chunk and print the corresponding content; # 2. If the current line is not empty and starts with data:, it indicates the start of a data chunk transmission, we remove the data: prefix and first check if it is the end symbol [DONE], if not, save the data content to the data variable; # 3. If the current line is not empty but does not start with data:, it indicates that the current line still belongs to the previous data chunk being transmitted, we append the content of the current line to the end of the data variable; if len(line) == 0: chunk = json.loads(data) # The processing logic here can be replaced with your business logic, printing is just to demonstrate the process choice = chunk["choices"][0] usage = choice.get("usage") if usage: print("total_tokens:", usage["total_tokens"]) delta = choice["delta"] role = delta.get("role") if role: print("role:", role) content = delta.get("content") if content: print(content, end="") data = "" # Reset data elif line.startswith("data: "): data = line.lstrip("data: ") # When the data chunk content is [DONE], it indicates that all data chunks have been sent, and the network connection can be disconnected if data == "[DONE]": break else: data = data + "\n" + line # We still add a newline character when appending content, as this data chunk may intentionally format the data in separate lines ``` ```js theme={null} const axios = require('axios'); // Use the axios library to make HTTP requests let data = { "model": "kimi-k3", "messages": [ // Specific messages ], "stream": true, }; // Use axios to send a chat request to the Kimi large language model and get the response r axios.post("https://api.moonshot.ai/v1/chat/completions", data, { responseType: 'stream' }).then(response => { let data = ''; response.data.on('data', chunk => { // Remove leading and trailing spaces from each line to better handle data chunks let line = chunk.toString().trim(); // Next, we need to handle three different cases: // 1. If the current line is empty, it indicates that the previous data chunk has been received (as mentioned earlier, the data chunk transmission ends with two newline characters), we can deserialize the data chunk and print the corresponding content; // 2. If the current line is not empty and starts with data:, it indicates the start of a data chunk transmission, we remove the data: prefix and first check if it is the end symbol [DONE], if not, save the data content to the data variable; // 3. If the current line is not empty but does not start with data:, it indicates that the current line still belongs to the previous data chunk being transmitted, we append the content of the current line to the end of the data variable; if (line === '') { try { let chunk = JSON.parse(data); // The processing logic here can be replaced with your business logic, printing is just to demonstrate the process let choice = chunk.choices[0]; let usage = choice.usage; if (usage) { console.log("total_tokens:", usage.total_tokens); } let delta = choice.delta; let role = delta.role; if (role) { console.log("role:", role); } let content = delta.content; if (content) { console.log(content); } } catch (error) { console.error("Error parsing JSON:", error); } data = ''; // Reset data } else if (line.startsWith('data: ')) { data = line.substring(6); // When the data chunk content is [DONE], it indicates that all data chunks have been sent, and the network connection can be disconnected if (data === '[DONE]') { response.data.destroy(); } } else { data += '\n' + line; // We still add a newline character when appending content, as this data chunk may intentionally format the data in separate lines } }); }).catch(error => { console.error("Error in request:", error); }); ``` Whatever the language, the basic steps for handling streaming output are the same: 1. Send an HTTP request with the `stream` parameter set to `true` in the request body; 2. Check the `Content-Type` in the response `Headers` — `text/event-stream` means the response is a streaming output; 3. Read the response line by line and parse the data chunks (in JSON format), locating chunk boundaries via the `data: ` prefix and newline characters `\n`; 4. A chunk whose content is `[DONE]` marks the end of the transmission. ## Multiple Responses (`n` Parameter) Current models (`kimi-k3`, `kimi-k2.7-code`, `kimi-k2.6`) fix `n` at `1` and do not support returning multiple responses in a single request. Passing an `n` greater than 1 returns a 400 error (`invalid n: only 1 is allowed for this model`) for both streaming and non-streaming requests. See the [Model Parameter Reference](/docs/api/models-overview) for per-model parameter constraints. # Main Concepts Source: https://platform.kimi.ai/docs/introduction Learn the core Kimi API concepts behind models, prompts, tokens, context windows, streaming, tool calling, and multimodal input. ## Text and Multimodal Models `kimi-k3` is Kimi's flagship model, built for long-horizon coding and end-to-end knowledge work, with native visual understanding; `kimi-k2.6` supports text, image, and video input, as well as thinking and non-thinking modes, and is suitable for conversation, code generation, visual understanding, and agent tasks. The input to a model is commonly called a "prompt", and clear instructions plus representative examples are the most effective way to get stable outputs. Other models are also available — see the [Model List](/docs/models) for details. ## Language Model Inference Service The language model inference service is an API service based on the pretrained models developed and trained by us (Moonshot AI). Today, the platform primarily exposes a Chat Completions interface for conversation, code generation, visual understanding, and agent tasks. Models do not directly access external resources such as the internet or databases by default, but you can extend them with official tools or custom tool calls when needed. ## Token Text generation models process text in units called Tokens. A Token represents a common sequence of characters. For example, a single English character like "antidisestablishmentarianism" might be broken down into a combination of several Tokens, while a short and common phrase like "word" might be represented by a single Token. Generally speaking, for a typical English text, 1 Token is roughly equivalent to 3-4 English characters. It is important to note that the total length of Input and Output cannot exceed the selected model's maximum context length. For example, `kimi-k3` supports a context window of up to 1M tokens. For other models' context lengths, see the [Model List](/docs/models). ## Rate Limits How do these rate limits work? Rate limits are measured in four ways: concurrency, RPM (requests per minute), TPM (tokens per minute), and TPD (tokens per day). The rate limit can be reached in any of these categories, depending on which one is hit first. For example, you might send 20 requests to ChatCompletions, each with only 100 Tokens, and you would hit the limit (if your RPM limit is 20), even if you haven't reached 200k Tokens in those 20 requests (assuming your TPM limit is 200k). For the gateway, for convenience, we calculate rate limits based on the max\_completion\_tokens parameter in the request. This means that if your request includes the max\_completion\_tokens parameter, we will use this parameter to calculate the rate limit. If your request does not include the max\_completion\_tokens parameter, we will use the default max\_completion\_tokens parameter to calculate the rate limit. After you make a request, we will determine whether you have reached the rate limit based on the number of Tokens in your request plus the number of max\_completion\_tokens in your parameter, regardless of the actual number of Tokens generated. In the billing process, we calculate the cost based on the number of Tokens in your request plus the actual number of Tokens generated. ### Other Important Notes: * Rate limits are enforced at the user level, not the key level. * Currently, we share rate limits across all models. ## Model List For all available models and their capabilities, see the [Model List](/docs/models) page. # Usage Guide ## Getting an API Key You need an API key to use our service. You can create an [API key](https://platform.kimi.ai/console/api-keys) in our [Console](https://platform.kimi.ai/console). ## Sending Requests You can use our Chat Completions API to send requests. You need to provide an API key and a model name. You can choose to use the default max\_completion\_tokens parameter or customize the max\_completion\_tokens parameter. You can refer to the [Chat API documentation](/docs/api/chat) for the calling method. ## Handling Responses Generally, we set a 2 hours timeout. If a single request exceeds this time, we will return a 504 error. If your request exceeds the rate limit, we will return a 429 error. If your request is successful, we will return a response in JSON format. If you need to quickly process some tasks, you can use the non-streaming mode of our Chat Completions API. In this mode, we will return all the generated text in one request. If you need more control, you can use the streaming mode. In this mode, we will return an [SSE](https://kimi.moonshot.cn/share/cr7boh3dqn37a5q9tds0) stream, where you can obtain the generated text. This can provide a better user experience, and you can also interrupt the request at any time without wasting resources. # Model List Source: https://platform.kimi.ai/docs/models Review currently available Kimi multimodal, coding, and Moonshot V1 models, plus migration guidance for discontinued models. Click [here](pricing/chat) to see more details of model price. Following the Kimi K3 launch, `kimi-k2.5` and the `moonshot-v1` series are no longer available to newly registered users (full platform sunset on August 31). Please switch to a newer model as soon as possible. ## Multi-modal Model | Model Name | Description | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `kimi-k3` | Kimi's most capable model to date, with 2.8 trillion parameters, native visual understanding, and a 1M-token context window, designed for frontier intelligence scenarios such as software engineering, knowledge work, and deep reasoning. | | `kimi-k2.7-code` | Kimi's dedicated coding model. It follows instructions more reliably in long contexts, completes coding tasks with higher success rates. Context 256k | | `kimi-k2.7-code-highspeed` | High-Speed version of Kimi K2.7 Code model, with output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience. | | `kimi-k2.6` | Supports both visual and text input, thinking and non-thinking modes, and dialogue and Agent tasks. Context 256k | | `kimi-k2.5` | Achieves open-source SoTA performance in Agent, code, visual understanding, and a range of general intelligent tasks. It also supports visual and text input, thinking and non-thinking modes, and dialogue and Agent tasks. Context 256k | ## Generation Model Moonshot V1 | Model Name | Description | | --------------------------------- | ----------------------------------------------------------------------------- | | `moonshot-v1-8k` | Suitable for generating short texts, context length 8k | | `moonshot-v1-32k` | Suitable for generating long texts, context length 32k | | `moonshot-v1-128k` | Suitable for generating very long texts, context length 128k | | `moonshot-v1-8k-vision-preview` | Vision model, understands image content and outputs text, context length 8k | | `moonshot-v1-32k-vision-preview` | Vision model, understands image content and outputs text, context length 32k | | `moonshot-v1-128k-vision-preview` | Vision model, understands image content and outputs text, context length 128k | > Note: The only difference between these Moonshot V1 models is their maximum context length (including input and output), there is no difference in effect. ## Deprecated Models > The `kimi-k2` series models were officially discontinued on **May 25, 2026** and are no longer maintained or supported. Please use the latest Kimi model [kimi-k3](/docs/guide/kimi-k3-quickstart) for continued support and enhanced reasoning capabilities. | Model Name | Description | | ------------------------ | ----------- | | `kimi-k2-0905-preview` | Deprecated | | `kimi-k2-0711-preview` | Deprecated | | `kimi-k2-turbo-preview` | Deprecated | | `kimi-k2-thinking` | Deprecated | | `kimi-k2-thinking-turbo` | Deprecated | > `kimi-latest` was officially discontinued on **January 28, 2026** and is no longer maintained or supported. Please use the latest Kimi model [kimi-k3](/docs/guide/kimi-k3-quickstart) for continued support and enhanced reasoning capabilities. > `kimi-thinking-preview` was officially discontinued on **November 11, 2025** and is no longer maintained or supported. We recommend upgrading to the latest model [kimi-k3](/docs/guide/kimi-k3-quickstart) for continued support and enhanced reasoning capabilities. For further assistance, please [contact sales](https://platform.kimi.ai/contact-sales). # Quickstart Source: https://platform.kimi.ai/docs/overview Create an API Key and complete your first Kimi API call. Kimi API lets you interact with Kimi models and is compatible with the OpenAI API format. Prepare an API Key, choose a model, and configure `base_url` to make requests through the HTTP API, Python SDK, or Node.js SDK. [Kimi K3](/docs/guide/kimi-k3-quickstart) is now officially available — Kimi's most capable model to date, with a 1M-token context window and native visual understanding. It is well suited for programming agent scenarios such as [Claude Code](guide/claude-code-kimi), as well as knowledge work and deep reasoning. ## Get started Visit the [Kimi API Platform](https://platform.kimi.ai/), sign in, then go to [API Keys](https://platform.kimi.ai/console/api-keys) to create and copy an API Key. Keep your API Key secure. Do not share it with others or hard-code it directly in your application. We recommend storing it in an environment variable: ```bash theme={null} export MOONSHOT_API_KEY="YOUR_KIMI_API_KEY" ``` Sign in to the platform to access the console, development workspace, and user center. Create, copy, and manage the keys used for API calls. For a quickstart, we recommend starting with Kimi K3; you can also choose Kimi K2.7 Code or Kimi K2.6 for specific scenarios. Kimi K3 is our flagship model for long-horizon coding and end-to-end knowledge work — 2.8 trillion parameters, a 1M-token context window, and industry-leading intelligence. A coding-focused model with a 256K context window, text/image/video input, and thinking mode. Choose `kimi-k2.7-code-highspeed` when you need higher output speed. A powerful general-purpose model with a 256K context window, text/image/video input, and both thinking and non-thinking modes. Use it for general chat, agent tasks, visual understanding, and complex reasoning. If you are not sure which model to choose, start with `kimi-k3`. If your task is mainly code generation, code editing, or programming agents and you need higher output speed, choose `kimi-k2.7-code-highspeed`. Kimi API is compatible with the OpenAI API format, so you can choose the integration method that best fits your technology stack. A standard REST API for any language or custom server-side integration. Test prompts, model behavior, and business examples without writing code. The following examples use the Kimi K3 model. Replace `MOONSHOT_API_KEY` with the API Key you created on the platform, or set an environment variable with the same name before running the examples. The examples on this page use the latest model `kimi-k3` by default. K3 configures reasoning effort with the top-level `reasoning_effort` request field (supports `"low"` / `"high"` / `"max"`, default `"max"`). To use another model such as `kimi-k2.6` or `kimi-k2.5`, just replace the `model` field — parameter configurations differ across models. See the [Model Parameter Reference](/docs/api/models-overview). To use the high-speed model for coding scenarios, replace `kimi-k3` in the examples with `kimi-k2.7-code-highspeed`; to call Kimi K2.6, replace it with `kimi-k2.6`. ```python theme={null} import os from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) completion = client.chat.completions.create( model="kimi-k3", messages=[ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are especially good at conversations in Chinese and English. You provide users with safe, helpful, and accurate answers. You also refuse to answer any questions involving terrorism, racism, pornography, violence, or similar harmful content. Moonshot AI is a proper noun and must not be translated into other languages."}, {"role": "user", "content": "Hi, my name is Li Lei. What is 1+1?"} ] ) print(completion.choices[0].message.content) ``` ```bash theme={null} curl https://api.moonshot.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $MOONSHOT_API_KEY" \ -d '{ "model": "kimi-k3", "messages": [ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are especially good at conversations in Chinese and English. You provide users with safe, helpful, and accurate answers. You also refuse to answer any questions involving terrorism, racism, pornography, violence, or similar harmful content. Moonshot AI is a proper noun and must not be translated into other languages."}, {"role": "user", "content": "Hi, my name is Li Lei. What is 1+1?"} ] }' ``` ```js theme={null} const OpenAI = require("openai"); const client = new OpenAI({ apiKey: process.env.MOONSHOT_API_KEY, baseURL: "https://api.moonshot.ai/v1", }); async function main() { const completion = await client.chat.completions.create({ model: "kimi-k3", messages: [ {"role": "system", "content": "You are Kimi, an AI assistant provided by Moonshot AI. You are especially good at conversations in Chinese and English. You provide users with safe, helpful, and accurate answers. You also refuse to answer any questions involving terrorism, racism, pornography, violence, or similar harmful content. Moonshot AI is a proper noun and must not be translated into other languages."}, {"role": "user", "content": "Hi, my name is Li Lei. What is 1+1?"} ] }); console.log(completion.choices[0].message.content); } main(); ``` Before running the examples, prepare: 1. Python 3.8 or later, or Node.js 18 or later. 2. OpenAI SDK 1.0.0 or later. Kimi API is compatible with the OpenAI API format, so you can call Kimi directly with the OpenAI Python or Node.js SDK. ```text theme={null} pip install --upgrade 'openai>=1.0' # Python npm install openai@latest # Node.js ``` 3. An API Key. Create an API Key in the [Kimi API Platform](https://platform.kimi.ai/console/api-keys), then pass it to the `OpenAI Client` so the platform can identify your account. If the code runs successfully without errors, you will see output similar to: ```text theme={null} Hello, Li Lei! 1+1 equals 2. This is a basic arithmetic question. If you have any other questions or need help, feel free to let me know. ``` *Note: Because Kimi models are nondeterministic, the actual response may not exactly match the example above.* ## Explore more features Enable `stream` to receive tokens as they are generated. Useful for chat, code generation, and long-form text output. Maintain a `messages` list to preserve context history, letting the model remember the conversation. Kimi K3, Kimi K2.7 Code, and Kimi K2.6 all support text, image, and video input. Let the model call external functions or APIs for Agent tasks, web search, and complex workflows. Force the model to output valid JSON for structured data extraction and downstream integration. Use reasoning capabilities for complex tasks, multi-step tool use, and Agent workflows. ### Streaming output ```json theme={null} { "model": "kimi-k3", "messages": [ { "role": "user", "content": "Please explain what recursion is and give a Python example." } ], "stream": true } ``` ### Multimodal input ```json theme={null} { "model": "kimi-k3", "messages": [ { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "data:image/png;base64,..." } }, { "type": "text", "text": "Please describe this image." } ] } ] } ``` For larger videos, or images and videos that need to be referenced multiple times, we recommend using file uploads. Images should be no larger than 4K resolution, and videos should be no larger than 1080p. ## Related resources View Kimi K3 capabilities, calling examples, and best practices. View Kimi K2.7 Code capabilities, calling examples, and best practices. View Kimi K2.6 capabilities, image/video understanding examples, and tool calling guidance. View currently available model names and descriptions. Ask questions, share feedback, and showcase what you are building with Kimi. # Platform Changelog Source: https://platform.kimi.ai/docs/platform-changelog Review historical Kimi Open Platform feature releases, model launches, product improvements, and issue fixes. This page is updated periodically with Kimi Open Platform product updates and related documentation changes. ## April 7, 2025 * Reduced model product pricing. * Added support for inviting and managing organization members. * Fixed an issue where the cursor could not move in the name field when creating a project. ## February 17, 2025 * Launched the `kimi-latest` model. * Added support for exporting organization monthly bills. * Fixed project rate limit display issues. * Added support for project daily and monthly spending alerts. ## January 13, 2025 * Launched the `moonshot-v1-vision-preview` model. * Added support for organization project management. * Restored WeChat Pay QR code payments. * Added support for overseas phone number registration and login. ## December 2, 2024 * Optimized the resource management list copy interaction to use hover and click. * Optimized resource list sorting by upload time, from newest to oldest. * Added support for multiple accounts under one verified business entity. * Fixed an invoice cancellation failure issue. ## November 4, 2024 * Context Caching is now available to all users. * Cache renewal no longer charges the creation fee. * Added terms and agreements content to the documentation center. * Updated and optimized API Key copy. * Fixed frontend flickering after successful payment. * Added a retry mechanism for invoice cancellation failures. * Fixed issues with editing and deleting API Keys. ## September 30, 2024 * Added frontend support for file resource management. * Split phone number rebinding into two frontend verification steps for the old and new phone numbers. * Added verification-code purpose descriptions to SMS messages. * Fixed an issue where the invoiceable amount was displayed incorrectly. * Fixed an invoice issuance failure caused by spaces in the tax identification number. * Added bank transfer processing time display for business verification. * Added documentation for automatic disconnection and reconnection handling. * Launched web search. ## August 28, 2024 * Launched `moonshot-v1-auto`. * Added support for custom account balance alerts. * Added support for phone number rebinding. * Added support for account password login. * Reduced Cache storage costs. * Published the MoonPalace user guide. * Released the Kimi Enterprise API. * Added tier level display to basic user information. ## July 31, 2024 * Released MoonPalace, the Kimi API debugging tool. * Optimized pagination for Context Caching management. * Launched user spending analysis. * Added an embedded Context Caching calculator entry point in the documentation. * Relaxed the company name length check for business verification. * Added support for changing an individual-verified user into a business-verified user. * Updated the API getting started guide. ## July 10, 2024 * Opened the Context Caching public beta to Tier 3 through Tier 5 users. * Added the developer community QR code. * Published the third Context Caching practice blog for Kimi API Assistant. * Published the second Context Caching practice blog for Kimi API Assistant. * Published the blog "How Context Caching Saves Up to 90% of Invocation Costs for Kimi API Assistant". ## July 1, 2024 * Officially launched the Context Caching public beta. ## June 28, 2024 * Added the WeCom customer service QR code. * Optimized the API Key count limit. * Published the first Context Caching practice blog for Kimi API Assistant. * Published the "Affordable Long-Text Processing" blog. * Added support for voucher validity periods. ## May 29, 2024 * Launched the Blog space. * Added Dark Mode support for Open Platform. * Launched invoice management. * Optimized WeChat Pay and Alipay QR code payments. ## April 30, 2024 * Launched Tool Calling. * Launched identity verification. * Launched corporate bank transfer. * Launched WeChat Pay and Alipay payments. * Added support for the balance monitoring API. # BatchJob Pricing Source: https://platform.kimi.ai/docs/pricing/batch Review Kimi BatchJob pricing for input, output, and cache-hit tokens, along with billing notes. ## Product Pricing **Explanation: Prices exclude applicable taxes. Specific tax obligations are subject to local tax regulations and will be calculated at checkout based on your jurisdiction.** Batch API inference costs are **60%** of the standard model price, ideal for large-scale tasks with low real-time requirements. Here, 1M = 1,000,000. The prices in the table represent the cost per 1M tokens consumed. ## Notes * Batch API supports `kimi-k2.7-code`, `kimi-k2.6` and `kimi-k2.5` models * Batch API is not subject to real-time concurrency limits, ideal for bulk tasks * Tasks must complete within the specified `completion_window`, otherwise they expire * See the [Batch API Guide](/docs/guide/use-batch-api) for detailed usage instructions # Model Inference Pricing Explanation Source: https://platform.kimi.ai/docs/pricing/chat Understand token billing, input and output charges, cache discounts, and pricing links for Kimi model inference. ## Concepts ### Billing Unit Token: A token represents a common sequence of characters. The number of tokens used for each English character may vary. For example, a single character like "antidisestablishmentarianism" might be broken down into several tokens, while a short and common phrase like "word" might use just one token. Generally speaking, for a typical English text, 1 token is roughly equivalent to 3-4 English characters. The exact number of tokens generated by each call can be obtained through the [Token Calculation API](/docs/api/estimate). #### Billing Logic Chat Completion API charges: We bill both the Input and Output based on usage. If you upload and extract content from a document and then pass the extracted content as Input to the model, the document content will also be billed based on usage. File-related interfaces (file content extraction/file storage) are **temporarily free**. In other words, if you only upload and extract a document, this API itself will not incur any charges. ## Model Pricing See detailed pricing for each model: Flagship model with a 1M-token context window Kimi's dedicated Coding model, multi-modal model Supports visual and text input Classic generation model series; full platform sunset expected on August 31 # Multi-modal Model Kimi K2.5 Pricing Source: https://platform.kimi.ai/docs/pricing/chat-k25 Review Kimi K2.5 multimodal model pricing for input, output, and cache-hit tokens, along with billing notes. ## Product Pricing **Explanation: Prices exclude applicable taxes. Specific tax obligations are subject to local tax regulations and will be calculated at checkout based on your jurisdiction.** Here, 1M = 1,000,000. The prices in the table represent the cost per 1M tokens consumed. ## Model Description The web search (`web_search`) is currently being updated. We do not recommend using this functionality in the near term. This documentation is outdated; please follow subsequent content updates. * Kimi K2.5 supports text, image, and video input, thinking and non-thinking modes, and dialogue and agent tasks. * Context length 256k, supports long thinking and deep reasoning. * Supports automatic context caching functionality, [ToolCalls](/docs/guide/use-kimi-api-to-complete-tool-calls), [JSON Mode](/docs/guide/use-json-mode-feature-of-kimi-api), [Partial Mode](/docs/guide/use-partial-mode-feature-of-kimi-api), and [internet search functionality](/docs/guide/use-web-search). # Kimi K2.6 Model Pricing Source: https://platform.kimi.ai/docs/pricing/chat-k26 Review Kimi K2.6 pricing for input, output, and cache-hit tokens, along with billing notes. ## Product Pricing **Explanation: Prices exclude applicable taxes. Specific tax obligations are subject to local tax regulations and will be calculated at checkout based on your jurisdiction.** Here, 1M = 1,000,000. The prices in the table represent the cost per 1M tokens consumed. ## Model Description The web search (`web_search`) is currently being updated. We do not recommend using this functionality in the near term. This documentation is outdated; please follow subsequent content updates. * Kimi K2.6 is a general-purpose model with stable long-horizon coding, instruction-following, and self-correction capabilities. It supports text, image, and video input, thinking and non-thinking modes, and dialogue and Agent tasks. * Context length 256k, supports long thinking and deep reasoning. * Supports automatic context caching functionality, [ToolCalls](/docs/guide/use-kimi-api-to-complete-tool-calls), [JSON Mode](/docs/guide/use-json-mode-feature-of-kimi-api), [Partial Mode](/docs/guide/use-partial-mode-feature-of-kimi-api), and [internet search functionality](/docs/guide/use-web-search). # Coding Model Kimi K2.7 Code Pricing Source: https://platform.kimi.ai/docs/pricing/chat-k27-code Review Kimi K2.7 Code and high-speed model pricing for input, output, and cache-hit tokens, along with billing notes. ## Product Pricing **Explanation: Prices exclude applicable taxes. Specific tax obligations are subject to local tax regulations and will be calculated at checkout based on your jurisdiction.** Here, 1M = 1,000,000. The prices in the table represent the cost per 1M tokens consumed. ## Model Description * Kimi K2.7 Code is a coding-focused model that completes programming tasks with higher success rates in long contexts. It supports text, image, and video input, thinking mode, dialogue, and agent tasks. * Kimi K2.7 Code HighSpeed is the high-speed version of Kimi K2.7 Code, the same model as Kimi K2.7 Code, but with an output speed of approximately 180 Tokens/s and up to 260 Tokens/s in short context scenarios, delivering a more extreme coding experience. * Context length 256k, supports long thinking and deep reasoning. * Supports automatic context caching functionality, [ToolCalls](/docs/guide/use-kimi-api-to-complete-tool-calls), [JSON Mode](/docs/guide/use-json-mode-feature-of-kimi-api), [Partial Mode](/docs/guide/use-partial-mode-feature-of-kimi-api). # Flagship Model Kimi K3 Pricing Source: https://platform.kimi.ai/docs/pricing/chat-k3 Review Kimi K3 flagship model pricing for input, output, and cache-hit tokens, along with billing notes. ## Product Pricing **Explanation: Prices exclude applicable taxes. Specific tax obligations are subject to local tax regulations and will be calculated at checkout based on your jurisdiction.** Here, 1M = 1,000,000. The prices in the table represent the cost per 1M tokens consumed. ## Model Description The web search (`web_search`) is currently being updated. We do not recommend using this functionality in the near term. This documentation is outdated; please follow subsequent content updates. * Kimi K3 is Kimi's flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window and industry-leading intelligence. See [Introducing Kimi K3](/docs/guide/kimi-k3-quickstart). * Always reasons and supports configuring its reasoning effort with the top-level `reasoning_effort` request field (`low` / `high` / `max`, default `max`). See [Reasoning Effort](/docs/guide/use-reasoning-effort). * Supports [automatic context caching](/docs/guide/use-context-caching-feature-of-kimi-api), [ToolCalls](/docs/guide/use-kimi-api-to-complete-tool-calls), [JSON Mode](/docs/guide/use-json-mode-feature-of-kimi-api), [structured output (`response_format` / JSON Schema)](/docs/guide/response_format), [Partial Mode](/docs/guide/use-partial-mode-feature-of-kimi-api), [internet search](/docs/guide/use-web-search), and more * New API capabilities in K3: [tool choice constraints (`tool_choice`)](/docs/guide/use-tool-choice) and [dynamically loaded tools](/docs/guide/use-dynamic-tool-loading) — see [Kimi K3 API Tool Calling Best Practices](/docs/guide/kimi-k3-tool-calling-best-practice) for combined usage # Generation Model Moonshot V1 Pricing Source: https://platform.kimi.ai/docs/pricing/chat-v1 Review Moonshot V1 generation and vision model pricing for input, output, and cache-hit tokens, along with billing notes. ## Product Pricing **Explanation: Prices exclude applicable taxes. Specific tax obligations are subject to local tax regulations and will be calculated at checkout based on your jurisdiction.** Here, 1M = 1,000,000. The prices in the table represent the cost per 1M tokens consumed. # Recharge and Rate Limiting Source: https://platform.kimi.ai/docs/pricing/limits Review Kimi Open Platform recharge requirements, account tiers, RPM, TPM, and TPD limits, and options for requesting higher capacity. Dear Kimi users: Due to a recent increase in high-frequency abnormal requests on the platform, which has affected the stability of cluster services, we plan to update the “Top-up Tiers and Rate Limits” rules in August. Please follow this page for updates. To ensure fair distribution of resources and prevent malicious attacks, we currently apply rate limits based on the cumulative recharge amount of each account. The specific limits are shown in the table below. If you have higher requirements, please contact us via email at [api-service@moonshot.ai](mailto:api-service@moonshot.ai). * To prevent abuse, you need to recharge at least \$1 to start using, and when your cumulative recharge reaches \$5, you will receive a \$5 voucher. ## Explanation of Rate Limits Concepts * Concurrency: The maximum number of requests from you that we can process at the same time. * RPM: Requests per minute, which means the maximum number of requests you can send to us in one minute. * TPM: Tokens per minute, which means the maximum number of tokens you can interact with us in one minute. * TPD: Tokens per day, which means the maximum number of tokens you can interact with us in one day. For more details, please refer to the [Rate Limits](/docs/introduction#rate-limits) section. ## Why Do We Implement Rate Limits? Rate limits are a common practice for API interfaces, and there are several reasons for it: * They help prevent abuse or misuse of the API. For example, malicious actors might try to overwhelm the API with a large number of requests, attempting to overload it or cause service disruptions. By setting rate limits, we can guard against such behavior. * Rate limits ensure fair access to the API for everyone. If one person or organization sends too many requests, it could slow down the API for everyone else. By limiting the number of requests a single user can send, we ensure that as many people as possible can use the API without experiencing slowdowns. * Rate limits help us manage the overall load on our cluster. A sudden surge in requests to the API could put pressure on the servers and lead to performance issues. By setting rate limits, we can maintain a smooth and consistent experience for all users. ## Special Notes * We will do our best to ensure normal usage for users, but when the cluster load reaches its capacity limit, we may take temporary measures to adjust the rate limits. * Vouchers do not count towards the cumulative recharge total. * When the system detects abnormal activity on an account, a risk-control rate-limiting policy is triggered. Once triggered, the restriction cannot be lifted. # WebSearch Pricing Source: https://platform.kimi.ai/docs/pricing/tools Review pricing, billing units, and usage notes for the Kimi web-search tool. ## Product Pricing **Explanation: Prices exclude applicable taxes. Specific tax obligations are subject to local tax regulations and will be calculated at checkout based on your jurisdiction.** ## Internet Search Billing Logic When you add the `$web_search` tool in `tools` and receive a response with `finish_reason = tool_calls` and `tool_call.function.name = $web_search`, we charge a fee of \$0.005 for the `$web_search` call. If the response has `finish_reason = stop`, no call fee will be charged. Additionally, when using `$web_search`, we still charge for the Tokens generated by the `/chat/completions` interface based on the model size. **It is important to note that when the `$web_search` tool is triggered, the search results are also counted in the Tokens. The number of Tokens occupied by the search results can be obtained from the returned `tool_call.function.arguments`.** For example, if the content of the `$web_search` occupies 4k Tokens, these 4k Tokens will be included in the total Tokens when the caller makes the next call to the `/chat/completions` interface. The total billing Tokens will be: ```text theme={null} total_tokens = prompt_tokens + search_tokens + completions_tokens ``` *Note: If you stop after triggering the `$web_search` without continuing with `tool_calls`, we will only charge the tool call fee of \$0.005, and the Tokens occupied by the search content will not be billed.*