48 terms
Definitions in this topic
- LLM foundationsAI21 LabsAI21 Labs explained as a language model developer and enterprise AI platform provider.
- LLM foundationsAlibaba Cloud AIAlibaba Cloud AI explained as a cloud platform for foundation models, model development, and hosted inference.
- ProductionAmazon BedrockAmazon Bedrock explained as AWS's managed platform for accessing and operating multiple foundation models.
- LLM foundationsAnthropicAnthropic explained as the company behind Claude, its developer platform, and the controls applications still need.
- ProductionAzure AI FoundryAzure AI Foundry, now Microsoft Foundry, explained as Microsoft's platform for building and operating AI applications.
- ProductionBring Your Own ModelBring your own model defined, including deployment patterns, benefits, and operating responsibilities.
- ProductionCerebrasCerebras explained as an AI compute and inference provider using wafer-scale processor architecture.
- LLM foundationsCohereCohere explained as an enterprise AI provider for generation, retrieval, reranking, and multilingual systems.
- LLM foundationsDeepSeekDeepSeek explained as an AI model developer, including open models, hosted access, and production evaluation criteria.
- LLM foundationsFine-tuningFine-tuning explained: what it actually changes, and why most teams should try RAG or better context first.
- ProductionFireworks AIFireworks AI explained as managed infrastructure for serving, customizing, and scaling generative models.
- LLM foundationsGoogle DeepMindGoogle DeepMind explained: its research role, relationship with Google, and how businesses access its models.
- ProductionGoogle Vertex AIGoogle Vertex AI explained as Google Cloud's managed platform for models, agents, evaluation, and AI operations.
- ProductionGroqGroq explained as an AI inference provider built around LPU hardware, and how it differs from xAI's Grok.
- ProductionHosted InferenceHosted inference defined, including its operational benefits, tradeoffs, and production evaluation criteria.
- ProductionHugging FaceHugging Face explained as the AI platform for models, datasets, open-source libraries, and hosted inference.
- LLM foundationsInferenceInference explained: what actually happens per API call, why it costs money every time, and how to keep it fast.
- ProductionInference ProviderInference provider defined, including how managed model serving differs from model development.
- LLM foundationsLLMWhat an LLM actually is from a buyer's seat: what it can carry in production, and where it needs a system built around it.
- ProductionLLM ProviderLLM provider defined, including model developers, hosting platforms, and the criteria that matter in production.
- LLM foundationsMeta AIMeta AI explained as a model developer, including open weights, hosting choices, and deployment responsibility.
- LLM foundationsMiniMaxMiniMax explained as a multimodal AI developer providing models and services across text, audio, image, and video.
- LLM foundationsMistral AIMistral AI explained as a European model provider offering hosted APIs and selected open weight models.
- LLM foundationsMixture of ExpertsMixture of Experts explained: how sparse routing gives language models more capacity without running every parameter.
- ProductionModel distillationModel distillation explained: how a small student learns from a big teacher, and when it beats fine-tuning or quantization.
- ProductionModel routingModel routing explained: choosing an LLM per request to balance capability, latency, cost, privacy, and availability.
- LLM foundationsMoonshot AIMoonshot AI explained as the developer of Kimi models, with practical criteria for production evaluation.
- ProductionMulti-Provider AIMulti-provider AI defined, including routing, resilience, portability, and the operational cost of multiple vendors.
- LLM foundationsMultimodal LLMMultimodal LLMs explained: models that combine language reasoning with images, audio, video, or other data types.
- ProductionNVIDIA AINVIDIA AI explained across GPUs, inference software, models, and enterprise deployment infrastructure.
- LLM foundationsOpen-weight modelOpen-weight models explained: downloadable parameters, licenses, and the open-source distinction.
- LLM foundationsOpenAIOpenAI explained as an AI model provider, including its APIs, platform role, and production considerations.
- ProductionOpenRouterOpenRouter explained as a shared API for accessing, comparing, and routing requests across many AI models.
- ProductionQuantizationQuantization explained: what fewer bits per weight buy you, what quality they cost, and when a compressed model is the right call.
- LLM foundationsReasoning modelsReasoning models explained: what extended thinking actually buys you, and when the extra latency is worth paying.
- ProductionReplicateReplicate explained as a hosted API platform for running community and publisher-provided AI models.
- ProductionServerless InferenceServerless inference defined, including autoscaling benefits and the latency, capacity, and cost tradeoffs.
- LLM foundationsSmall Language ModelSmall language model explained: when a compact model is faster, cheaper, and more private than a general LLM.
- LLM foundationsSpeculative DecodingSpeculative decoding explained: how draft and target models verify several tokens at once to reduce LLM latency.
- LLM foundationsTemperatureTemperature explained: how the sampling parameter trades determinism for variety, and how to set it per task.
- LLM foundationsTest-Time ComputeTest-time compute explained: spending more inference budget on reasoning, sampling, verification, and answer selection.
- ProductionTogether AITogether AI explained as an infrastructure provider for hosted inference, training, and deployment of open models.
- LLM foundationsTokenToken explained: what a token actually is, why it is not a word, and where the count quietly adds up.
- LLM foundationsTokenizationTokenization explained: how language models convert text into tokens and why token counts affect context, cost, and output.
- LLM foundationsTransformer architectureTransformer architecture explained: how attention connects tokens and powers modern language models.
- LLM foundationsVision-language modelVision-language models explained: combining visual inputs with language instructions for analysis, extraction, and agents.
- LLM foundationsxAIxAI explained as the company behind Grok, including API access and what businesses should evaluate.
- LLM foundationsZ.aiZ.ai explained as a provider of GLM foundation models and developer services for AI applications.