You're reading from Generative AI for Cloud Solutions Architect modern AI LLMs in secure, scalable, and ethical cloud environments

Product type Paperback

Published in Apr 2024

Publisher Packt

ISBN-13 9781835084786

Length 300 pages

Edition 1st Edition

Languages

Python

Tools

ChatGPT

Concepts

GPT/LLMs

Authors (2):

Paul Singh

Anurag Karuparti

View More author details

Table of Contents (18) Chapters

Preface

1. Part 1:Integrating Cloud Power with Language Breakthroughs FREE CHAPTER

2. Chapter 1: Cloud Computing Meets Generative AI: Bridging Infinite Impossibilities

3. Chapter 2: NLP Evolution and Transformers: Exploring NLPs and LLMs

4. Part 2: Techniques for Tailoring LLMs

5. Chapter 3: Fine-Tuning – Building Domain-Specific LLM Applications

6. Chapter 4: RAGs to Riches: Elevating AI with External Data

7. Chapter 5: Effective Prompt Engineering Techniques: Unlocking Wisdom Through AI

8. Part 3: Developing, Operationalizing, and Scaling Generative AI Applications

9. Chapter 6: Developing and Operationalizing LLM-based Apps: Exploring Dev Frameworks and LLMOps

10. Chapter 7: Deploying ChatGPT in the Cloud: Architecture Design and Scaling Strategies

11. Part 4: Building Safe and Secure AI – Security and Ethical Considerations

12. Chapter 8: Security and Privacy Considerations for Gen AI – Building Safe and Secure LLMs

13. Chapter 9: Responsible Development of AI Solutions: Building with Integrity and Care

14. Part 5: Generative AI – What’s Next?

15. Chapter 10: The Future of Generative AI – Trends and Emerging Use Cases

16. Index

Why subscribe?

17. Other Books You May Enjoy

Understanding limits

Any large-scale cloud deployment needs to be “enterprise-ready,” ensuring both the end user experience is acceptable and the business objectives and requirements are met. “Acceptable” is a loose term that can vary per user and workload. To understand how to scale to meet any user or business requirements, as the appetite for a service increases, we must first understand the basic limits, such as token limits. We covered these limits for most of the common generative AI GPT models in Chapter 5, however, we will quickly revisit them here.

As organizations scale up using an enterprise-ready service, such as Azure OpenAI, there are rate limits on how fast tokens are processed in the prompt+completion request. There is a limit to how many text prompts can be sent due to these token limits for each model that can be consumed in a single prompt+completion. It is important to note that the overall size of tokens for rate limiting includes...

The rest of the chapter is locked

You're reading from Generative AI for Cloud Solutions Architect modern AI LLMs in secure, scalable, and ethical cloud environments

Table of Contents (18) Chapters

Understanding limits

Unlock this book and the full library FREE for 7 days

Authors (2)

Personalised recommendations for you