Skip to content

bbstrader.models

Quant toolkit for signals and risk: NLP sentiment analysis / topic modeling, and portfolio optimization.

models

Overview

The Models Module provides a collection of quantitative models for financial analysis and decision-making. It includes tools for portfolio optimization and natural language processing (NLP) to extract insights from financial text data. This module is designed to support quantitative trading strategies by providing a robust framework for financial modeling.

Features

  • Portfolio Optimization: Implements techniques to optimize portfolio allocation, helping to maximize returns and manage risk.
  • Natural Language Processing (NLP): Provides tools for analyzing financial news and other text-based data to gauge market sentiment.
  • Extensible Design: Structured to allow for the easy addition of new quantitative models and algorithms.

Components

  • Optimization: Contains portfolio optimization models and related utilities.
  • NLP: Includes tools and models for natural language processing tailored for financial applications.

Examples

from bbstrader.models import optimized_weights

Assuming 'returns' is a DataFrame of asset returns

optimal_weights = optimized_weights(returns=returns) print(optimal_weights)

Notes

This module is focused on providing the analytical tools for quantitative analysis. The models can be integrated into trading strategies to provide data-driven signals.

SentimentAnalyzer

SentimentAnalyzer()

Bases: object

A financial sentiment analysis tool that processes and analyzes sentiment from news articles, social media posts, and financial reports.

This class utilizes NLP techniques to preprocess text and apply sentiment analysis using VADER (SentimentIntensityAnalyzer) and optional TextBlob for enhanced polarity scoring.

Initializes the SentimentAnalyzer class by downloading necessary NLTK resources and loading the SpaCy NLP model.

  • Downloads NLTK tokenization (punkt) and stopwords.
  • Loads the en_core_web_sm SpaCy model with Named Entity Recognition (NER) disabled.
  • Initializes VADER's SentimentIntensityAnalyzer for sentiment scoring.
Source code in src/bbstrader/models/nlp.py
def __init__(self):
    """
    Initializes the SentimentAnalyzer class by downloading necessary
    NLTK resources and loading the SpaCy NLP model.

    - Downloads NLTK tokenization (`punkt`) and stopwords.
    - Loads the `en_core_web_sm` SpaCy model with Named Entity Recognition (NER) disabled.
    - Initializes VADER's SentimentIntensityAnalyzer for sentiment scoring.

    """
    _require_nlp()
    nltk.download("punkt", quiet=True)
    nltk.download("punkt_tab", quiet=True)
    nltk.download("stopwords", quiet=True)

    self.analyzer = SentimentIntensityAnalyzer()
    self._stopwords = set(stopwords.words("english"))

    try:
        self.nlp = spacy.load("en_core_web_sm", disable=["ner"])
    except OSError:
        print("Downloading 'en_core_web_sm' model for spaCy...")
        download("en_core_web_sm")
        self.nlp = spacy.load("en_core_web_sm", disable=["ner"])

    self.news = FinancialNews()

preprocess_text

preprocess_text(text: str)

Preprocesses the input text by performing the following steps: 1. Converts text to lowercase. 2. Removes URLs. 3. Removes all non-alphabetic characters (punctuation, numbers, special symbols). 4. Tokenizes the text into words. 5. Removes stop words. 6. Lemmatizes the words using SpaCy, excluding pronouns.

Parameters:

Name Type Description Default
text str

The input text to preprocess.

required

Returns:

Name Type Description
str

The cleaned and lemmatized text.

Source code in src/bbstrader/models/nlp.py
def preprocess_text(self, text: str):
    """
    Preprocesses the input text by performing the following steps:
    1. Converts text to lowercase.
    2. Removes URLs.
    3. Removes all non-alphabetic characters (punctuation, numbers, special symbols).
    4. Tokenizes the text into words.
    5. Removes stop words.
    6. Lemmatizes the words using SpaCy, excluding pronouns.

    Args:
        text (str): The input text to preprocess.

    Returns:
        str: The cleaned and lemmatized text.
    """
    if not isinstance(text, str):
        raise ValueError(
            f"{self.__class__.__name__}: preprocess_text expects a string, got {type(text)}"
        )
    text = text.lower()
    text = re.sub(r"http\S+", "", text)
    text = re.sub(r"[^a-zA-Z\s]", "", text)

    words = word_tokenize(text)
    words = [word for word in words if word not in self._stopwords]

    doc = self.nlp(" ".join(words))
    words = [t.lemma_ for t in doc if t.lemma_ != "-PRON-"]

    return " ".join(words)

analyze_sentiment

analyze_sentiment(texts, lexicon=None, textblob=False) -> float

Analyzes the sentiment of a list of texts using VADER or TextBlob.

Steps: 1. If a custom lexicon is provided, updates the VADER lexicon. 2. If textblob is set to True, computes sentiment using TextBlob. 3. Otherwise, preprocesses the text and computes sentiment using VADER. 4. Returns the average sentiment score of all input texts.

Parameters:

Name Type Description Default
texts list of str

A list of text inputs to analyze.

required
lexicon dict

A custom sentiment lexicon to update VADER's default lexicon.

None
textblob bool

If True, uses TextBlob for sentiment analysis instead of VADER.

False

Returns:

Name Type Description
float float

The average sentiment score across all input texts. - Positive values indicate positive sentiment. - Negative values indicate negative sentiment. - Zero indicates neutral sentiment.

Source code in src/bbstrader/models/nlp.py
def analyze_sentiment(self, texts, lexicon=None, textblob=False) -> float:
    """
    Analyzes the sentiment of a list of texts using VADER or TextBlob.

    Steps:
    1. If a custom lexicon is provided, updates the VADER lexicon.
    2. If `textblob` is set to True, computes sentiment using TextBlob.
    3. Otherwise, preprocesses the text and computes sentiment using VADER.
    4. Returns the average sentiment score of all input texts.

    Args:
        texts (list of str): A list of text inputs to analyze.
        lexicon (dict, optional): A custom sentiment lexicon to update VADER's default lexicon.
        textblob (bool, optional): If True, uses TextBlob for sentiment analysis instead of VADER.

    Returns:
        float: The average sentiment score across all input texts.
               - Positive values indicate positive sentiment.
               - Negative values indicate negative sentiment.
               - Zero indicates neutral sentiment.
    """
    if lexicon is not None:
        self.analyzer.lexicon.update(lexicon)
    if textblob:
        blob = TextBlob(" ".join(texts))
        return blob.sentiment.polarity
    sentiment_scores = [
        self.analyzer.polarity_scores(self.preprocess_text(text))["compound"]
        for text in texts
    ]
    avg_sentiment = (
        sum(sentiment_scores) / len(sentiment_scores) if sentiment_scores else 0.0
    )
    return avg_sentiment

get_sentiment_for_tickers

get_sentiment_for_tickers(tickers: List[str] | List[Tuple[str, str]], lexicon=None, asset_type='stock', top_news=10, **kwargs) -> Dict[str, float]

Compute sentiment scores for a list of financial tickers based on news and social media data.

Process
  1. Collect news articles and posts related to each ticker from various sources:
  2. Yahoo Finance News
  3. Google Finance News
  4. Reddit posts
  5. Financial Modeling Prep (FMP) news
  6. Analyze sentiment from each source:
  7. Uses VADER for Yahoo and Google Finance news.
  8. Uses TextBlob for Reddit and FMP news.
  9. Compute an overall sentiment score using a weighted average approach.
Parameters

tickers : list of str or list of tuple A list of asset tickers to analyze. * If using tuples, the first element is the ticker and the second is the asset type. * If using a single string, the asset type must be specified or defaults to "stock". lexicon : dict, optional A custom sentiment lexicon to update VADER's default lexicon. Default is None. asset_type : str, optional The type of asset. Default is "stock". Supported types include: * "stock": Stock symbols (e.g., AAPL, MSFT) * "etf": Exchange-traded funds (e.g., SPY, QQQ) * "future": Futures contracts (e.g., CL=F for crude oil) * "forex": Forex pairs (e.g., EURUSD=X, USDJPY=X) * "crypto": Cryptocurrency pairs (e.g., BTC-USD, ETH-USD) * "index": Stock market indices (e.g., ^GSPC for S&P 500) top_news : int, optional Number of news articles/posts to fetch per source. Default is 10. **kwargs : dict Additional parameters for API authentication and data retrieval. Must include: * fmp_api (str): API key for Financial Modeling Prep. * client_id, client_secret, user_agent (str): Credentials for Reddit API.

Returns

dict of str to float A dictionary mapping each ticker to its overall sentiment score. * Positive values indicate positive sentiment. * Negative values indicate negative sentiment. * Zero indicates neutral sentiment.

Notes

Ticker names must follow Yahoo Finance conventions.

Source code in src/bbstrader/models/nlp.py
def get_sentiment_for_tickers(
    self,
    tickers: List[str] | List[Tuple[str, str]],
    lexicon=None,
    asset_type="stock",
    top_news=10,
    **kwargs,
) -> Dict[str, float]:
    """
    Compute sentiment scores for a list of financial tickers based on news and social media data.

    Process
    -------
    1. Collect news articles and posts related to each ticker from various sources:
    * Yahoo Finance News
    * Google Finance News
    * Reddit posts
    * Financial Modeling Prep (FMP) news
    2. Analyze sentiment from each source:
    * Uses VADER for Yahoo and Google Finance news.
    * Uses TextBlob for Reddit and FMP news.
    3. Compute an overall sentiment score using a weighted average approach.

    Parameters
    ----------
    tickers : list of str or list of tuple
        A list of asset tickers to analyze.
        * If using tuples, the first element is the ticker and the second is the asset type.
        * If using a single string, the asset type must be specified or defaults to "stock".
    lexicon : dict, optional
        A custom sentiment lexicon to update VADER's default lexicon. Default is None.
    asset_type : str, optional
        The type of asset. Default is "stock".
        Supported types include:
        * "stock": Stock symbols (e.g., AAPL, MSFT)
        * "etf": Exchange-traded funds (e.g., SPY, QQQ)
        * "future": Futures contracts (e.g., CL=F for crude oil)
        * "forex": Forex pairs (e.g., EURUSD=X, USDJPY=X)
        * "crypto": Cryptocurrency pairs (e.g., BTC-USD, ETH-USD)
        * "index": Stock market indices (e.g., ^GSPC for S&P 500)
    top_news : int, optional
        Number of news articles/posts to fetch per source. Default is 10.
    **kwargs : dict
        Additional parameters for API authentication and data retrieval. Must include:
        * fmp_api (str): API key for Financial Modeling Prep.
        * client_id, client_secret, user_agent (str): Credentials for Reddit API.

    Returns
    -------
    dict of str to float
        A dictionary mapping each ticker to its overall sentiment score.
        * Positive values indicate positive sentiment.
        * Negative values indicate negative sentiment.
        * Zero indicates neutral sentiment.

    Notes
    -----
    Ticker names must follow Yahoo Finance conventions.
    """

    sentiment_results = {}

    # Suppress stdout/stderr from underlying  libraries during execution
    with open(os.devnull, "w") as devnull:
        with (
            contextlib.redirect_stdout(devnull),
            contextlib.redirect_stderr(devnull),
        ):
            with ThreadPoolExecutor() as executor:
                # Map each future to its ticker for easy result lookup
                future_to_ticker = {}
                for ticker_info in tickers:
                    # Normalize input to (ticker, asset_type)
                    if isinstance(ticker_info, tuple):
                        ticker_symbol, ticker_asset_type = ticker_info
                    else:
                        ticker_symbol, ticker_asset_type = ticker_info, asset_type

                    if ticker_asset_type not in [
                        "stock",
                        "etf",
                        "future",
                        "forex",
                        "crypto",
                        "index",
                    ]:
                        raise ValueError(
                            f"Unsupported asset type '{ticker_asset_type}' for {ticker_symbol}."
                        )

                    # Submit the job to the thread pool
                    future = executor.submit(
                        self._get_sentiment_for_one_ticker,
                        ticker=ticker_symbol,
                        asset_type=ticker_asset_type,
                        lexicon=lexicon,
                        top_news=top_news,
                        **kwargs,
                    )
                    future_to_ticker[future] = ticker_symbol

                # Collect results as they are completed
                for future in as_completed(future_to_ticker):
                    ticker_symbol = future_to_ticker[future]
                    try:
                        sentiment_score = future.result()
                        sentiment_results[ticker_symbol] = sentiment_score
                    except Exception:
                        sentiment_results[ticker_symbol] = (
                            0.0  # Assign a neutral score on error
                        )

    return sentiment_results

get_topn_sentiments

get_topn_sentiments(sentiments, topn=10)

Retrieves the top and bottom N assets based on sentiment scores.

Parameters:

Name Type Description Default
sentiments dict

A dictionary mapping asset tickers to their sentiment scores.

required
topn int

The number of top and bottom assets to return. Defaults to 10.

10

Returns:

Name Type Description
tuple

A tuple containing two lists: - bottom (list of tuples): The topn assets with the lowest sentiment scores, sorted in ascending order. - top (list of tuples): The topn assets with the highest sentiment scores, sorted in descending order.

Source code in src/bbstrader/models/nlp.py
def get_topn_sentiments(self, sentiments, topn=10):
    """
    Retrieves the top and bottom N assets based on sentiment scores.

    Args:
        sentiments (dict): A dictionary mapping asset tickers to their sentiment scores.
        topn (int, optional): The number of top and bottom assets to return. Defaults to 10.

    Returns:
        tuple: A tuple containing two lists:
            - bottom (list of tuples): The `topn` assets with the lowest sentiment scores, sorted in ascending order.
            - top (list of tuples): The `topn` assets with the highest sentiment scores, sorted in descending order.
    """
    sorted_sentiments = sorted(sentiments.items(), key=lambda x: x[1])
    bottom = sorted_sentiments[:topn]
    top = sorted_sentiments[-topn:]
    return bottom, top

visualize_sentiments

visualize_sentiments(sentiment_dict, mode='bar', top_n=10)

Visualizes sentiment scores for financial assets using different chart types.

Visualization Modes: - "bar": Displays a bar chart of the top N assets by sentiment score. - "scatter": Displays a scatter plot of sentiment scores.

Parameters:

Name Type Description Default
sentiment_dict dict

A dictionary mapping asset tickers to their sentiment scores.

required
mode str

The type of visualization to generate. Options: "bar" (default), "scatter".

'bar'
top_n int

The number of top tickers to display in the bar chart. Only applicable when mode is "bar".

10

Returns:

Name Type Description
None

Displays the sentiment visualization.

Source code in src/bbstrader/models/nlp.py
def visualize_sentiments(self, sentiment_dict, mode="bar", top_n=10):
    """
    Visualizes sentiment scores for financial assets using different chart types.

    Visualization Modes:
    - "bar": Displays a bar chart of the top N assets by sentiment score.
    - "scatter": Displays a scatter plot of sentiment scores.

    Args:
        sentiment_dict (dict): A dictionary mapping asset tickers to their sentiment scores.
        mode (str, optional): The type of visualization to generate.
                              Options: "bar" (default), "scatter".
        top_n (int, optional): The number of top tickers to display in the bar chart.
                               Only applicable when mode is "bar".

    Returns:
        None: Displays the sentiment visualization.
    """
    if mode == "bar":
        self._sentiment_bar(sentiment_dict, top_n=top_n)
    elif mode == "scatter":
        self._sentiment_scatter(sentiment_dict)

markowitz_weights

markowitz_weights(prices=None, rfr=0.0, freq=252, min_vol=False)

Calculates optimal portfolio weights using Markowitz's mean-variance optimization (Max Sharpe Ratio or Min Volatility) with multiple solvers.

Parameters

prices : pd.DataFrame, optional Price data for assets, where rows represent time periods and columns represent assets. rfr : float, optional Risk-free rate (default is 0.0). freq : int, optional Frequency of the data, such as 252 for daily returns in a year (default is 252). min_vol : bool, optional If True, optimizes for minimum volatility instead of maximum Sharpe ratio (default is False).

Returns

dict Dictionary containing the optimal asset weights for maximizing the Sharpe ratio or minimizing volatility, normalized to sum to 1.

Notes

This function attempts to maximize the Sharpe ratio by iterating through various solvers ('SCS', 'ECOS', 'OSQP') from the PyPortfolioOpt library. If a solver fails, it proceeds to the next one. If none succeed, an error message is printed for each solver that fails.

This function is useful for portfolio with a small number of assets, as it may not scale well for large portfolios.

Raises

Exception If all solvers fail, each will print an exception error message during runtime.

Source code in src/bbstrader/models/optimization.py
def markowitz_weights(prices=None, rfr=0.0, freq=252, min_vol=False):
    """
    Calculates optimal portfolio weights using Markowitz's mean-variance optimization (Max Sharpe Ratio or Min Volatility) with multiple solvers.

    Parameters
    ----------
    prices : pd.DataFrame, optional
        Price data for assets, where rows represent time periods and columns represent assets.
    rfr : float, optional
        Risk-free rate (default is 0.0).
    freq : int, optional
        Frequency of the data, such as 252 for daily returns in a year (default is 252).
    min_vol : bool, optional
        If True, optimizes for minimum volatility instead of maximum Sharpe ratio (default is False).

    Returns
    -------
    dict
        Dictionary containing the optimal asset weights for maximizing the Sharpe ratio or minimizing volatility, normalized to sum to 1.

    Notes
    -----
    This function attempts to maximize the Sharpe ratio by iterating through various solvers ('SCS', 'ECOS', 'OSQP')
    from the PyPortfolioOpt library. If a solver fails, it proceeds to the next one. If none succeed, an error message
    is printed for each solver that fails.

    This function is useful for portfolio with a small number of assets, as it may not scale well for large portfolios.

    Raises
    ------
    Exception
        If all solvers fail, each will print an exception error message during runtime.
    """
    returns = expected_returns.mean_historical_return(prices, frequency=freq)
    cov = risk_models.sample_cov(prices, frequency=freq)

    # Try different solvers to maximize Sharpe ratio
    for solver in ["SCS", "ECOS", "OSQP"]:
        ef = EfficientFrontier(
            expected_returns=returns,
            cov_matrix=cov,
            weight_bounds=(0, 1),
            solver=solver,
        )
        try:
            if min_vol:
                ef.min_volatility()
            else:
                ef.max_sharpe(risk_free_rate=rfr)
            return _normalize_weights(ef.clean_weights())
        except Exception as e:
            print(f"Solver {solver} failed with error: {e}")
    # Default to equal weighted if all solvers fail
    return _normalize_weights(equal_weighted(prices=prices))

hierarchical_risk_parity

hierarchical_risk_parity(prices=None, returns=None, freq=252)

Computes asset weights using Hierarchical Risk Parity (HRP) for risk-averse portfolio allocation.

Parameters

prices : pd.DataFrame, optional Price data for assets; if provided, daily returns will be calculated. returns : pd.DataFrame, optional Daily returns for assets. One of prices or returns must be provided. freq : int, optional Number of days to consider in calculating portfolio weights (default is 252).

Returns

dict Optimized asset weights using the HRP method, with asset weights summing to 1.

Raises

ValueError If neither prices nor returns are provided.

Notes

Hierarchical Risk Parity is particularly useful for portfolios with a large number of assets, as it mitigates issues of multicollinearity and estimation errors in covariance matrices by using hierarchical clustering.

Source code in src/bbstrader/models/optimization.py
def hierarchical_risk_parity(prices=None, returns=None, freq=252):
    """
    Computes asset weights using Hierarchical Risk Parity (HRP) for risk-averse portfolio allocation.

    Parameters
    ----------
    prices : pd.DataFrame, optional
        Price data for assets; if provided, daily returns will be calculated.
    returns : pd.DataFrame, optional
        Daily returns for assets. One of `prices` or `returns` must be provided.
    freq : int, optional
        Number of days to consider in calculating portfolio weights (default is 252).

    Returns
    -------
    dict
        Optimized asset weights using the HRP method, with asset weights summing to 1.

    Raises
    ------
    ValueError
        If neither `prices` nor `returns` are provided.

    Notes
    -----
    Hierarchical Risk Parity is particularly useful for portfolios with a large number of assets,
    as it mitigates issues of multicollinearity and estimation errors in covariance matrices by
    using hierarchical clustering.
    """
    warnings.filterwarnings("ignore")
    if returns is None and prices is None:
        raise ValueError("Either prices or returns must be provided")
    if returns is None:
        returns = prices.pct_change().dropna(how="all")
    # Remove duplicate columns and index
    returns = returns.loc[:, ~returns.columns.duplicated()]
    returns = returns.loc[~returns.index.duplicated(keep="first")]
    hrp = HRPOpt(returns=returns.iloc[-freq:])
    return _normalize_weights(hrp.optimize())

equal_weighted

equal_weighted(prices=None, returns=None, round_digits=5)

Generates an equal-weighted portfolio by assigning an equal proportion to each asset.

Parameters

prices : pd.DataFrame, optional Price data for assets, where each column represents an asset. returns : pd.DataFrame, optional Return data for assets. One of prices or returns must be provided. round_digits : int, optional Number of decimal places to round each weight to (default is 5).

Returns

dict Dictionary with equal weights assigned to each asset, summing to 1.

Raises

ValueError If neither prices nor returns are provided.

Notes

Equal weighting is a simple allocation method that assumes equal importance across all assets, useful as a baseline model and when no strong views exist on asset return expectations or risk.

Source code in src/bbstrader/models/optimization.py
def equal_weighted(prices=None, returns=None, round_digits=5):
    """
    Generates an equal-weighted portfolio by assigning an equal proportion to each asset.

    Parameters
    ----------
    prices : pd.DataFrame, optional
        Price data for assets, where each column represents an asset.
    returns : pd.DataFrame, optional
        Return data for assets. One of `prices` or `returns` must be provided.
    round_digits : int, optional
        Number of decimal places to round each weight to (default is 5).

    Returns
    -------
    dict
        Dictionary with equal weights assigned to each asset, summing to 1.

    Raises
    ------
    ValueError
        If neither `prices` nor `returns` are provided.

    Notes
    -----
    Equal weighting is a simple allocation method that assumes equal importance across all assets,
    useful as a baseline model and when no strong views exist on asset return expectations or risk.
    """

    if returns is None and prices is None:
        raise ValueError("Either prices or returns must be provided")
    if returns is None:
        n = len(prices.columns)
        columns = prices.columns
    else:
        n = len(returns.columns)
        columns = returns.columns
    return {col: round(1 / n, round_digits) for col in columns}

black_litterman_weights

black_litterman_weights(prices=None, rfr=0.0, freq=252, views=None, view_confidences=None, pi=None, market_caps=None)

Computes portfolio weights using the Black-Litterman model.

Parameters

prices : pd.DataFrame Price data for assets. rfr : float, optional Risk-free rate (default is 0.0). freq : int, optional Frequency of the data (default is 252). views : dict, optional Investor's views on asset returns. view_confidences : list or np.array, optional Confidence levels for each view. pi : pd.Series, optional Market-implied prior returns. market_caps : pd.Series, optional Market capitalization of assets.

Returns

dict Optimal asset weights based on the Black-Litterman model.

Source code in src/bbstrader/models/optimization.py
def black_litterman_weights(
    prices=None,
    rfr=0.0,
    freq=252,
    views=None,
    view_confidences=None,
    pi=None,
    market_caps=None,
):
    """
    Computes portfolio weights using the Black-Litterman model.

    Parameters
    ----------
    prices : pd.DataFrame
        Price data for assets.
    rfr : float, optional
        Risk-free rate (default is 0.0).
    freq : int, optional
        Frequency of the data (default is 252).
    views : dict, optional
        Investor's views on asset returns.
    view_confidences : list or np.array, optional
        Confidence levels for each view.
    pi : pd.Series, optional
        Market-implied prior returns.
    market_caps : pd.Series, optional
        Market capitalization of assets.

    Returns
    -------
    dict
        Optimal asset weights based on the Black-Litterman model.
    """
    cov_matrix = risk_models.sample_cov(prices, frequency=freq)
    if pi is None:
        if market_caps is not None:
            # If market caps are provided, we can use them to compute the prior
            # This requires a benchmark, which we don't have here easily.
            # For simplicity, we use the mean historical return if pi is not provided.
            pi = expected_returns.mean_historical_return(prices, frequency=freq)
        else:
            pi = expected_returns.mean_historical_return(prices, frequency=freq)

    bl = BlackLittermanModel(
        cov_matrix,
        pi=pi,
        absolute_views=views,
        omega=None,
        view_confidences=view_confidences,
    )
    ret_bl = bl.bl_returns()
    ef = EfficientFrontier(ret_bl, cov_matrix)
    ef.max_sharpe(risk_free_rate=rfr)
    return _normalize_weights(ef.clean_weights())

optimized_weights

optimized_weights(prices=None, returns=None, rfr=0.0, freq=252, method='equal', **kwargs)

Selects an optimization method to calculate portfolio weights based on user preference.

Parameters

prices : pd.DataFrame, optional Price data for assets, required for certain methods. returns : pd.DataFrame, optional Returns data for assets, an alternative input for certain methods. freq : int, optional Number of days for calculating portfolio weights, such as 252 for a year's worth of daily returns (default is 252). method : str, optional Optimization method to use ('markowitz', 'hrp', or 'equal') (default is 'equal').

Returns

dict Dictionary containing optimized asset weights based on the chosen method.

Raises

ValueError If an unknown optimization method is specified.

Notes

This function integrates different optimization methods: - 'markowitz': mean-variance optimization with max Sharpe ratio - 'hrp': Hierarchical Risk Parity, for risk-based clustering of assets - 'equal': Equal weighting across all assets

Source code in src/bbstrader/models/optimization.py
def optimized_weights(
    prices=None, returns=None, rfr=0.0, freq=252, method="equal", **kwargs
):
    """
    Selects an optimization method to calculate portfolio weights based on user preference.

    Parameters
    ----------
    prices : pd.DataFrame, optional
        Price data for assets, required for certain methods.
    returns : pd.DataFrame, optional
        Returns data for assets, an alternative input for certain methods.
    freq : int, optional
        Number of days for calculating portfolio weights, such as 252 for a year's worth of daily returns (default is 252).
    method : str, optional
        Optimization method to use ('markowitz', 'hrp', or 'equal') (default is 'equal').

    Returns
    -------
    dict
        Dictionary containing optimized asset weights based on the chosen method.

    Raises
    ------
    ValueError
        If an unknown optimization method is specified.

    Notes
    -----
    This function integrates different optimization methods:
    - 'markowitz': mean-variance optimization with max Sharpe ratio
    - 'hrp': Hierarchical Risk Parity, for risk-based clustering of assets
    - 'equal': Equal weighting across all assets
    """
    if method == "markowitz":
        return markowitz_weights(prices=prices, rfr=rfr, freq=freq)
    elif method == "min_vol":
        return markowitz_weights(prices=prices, rfr=rfr, freq=freq, min_vol=True)
    elif method == "hrp":
        return hierarchical_risk_parity(prices=prices, returns=returns, freq=freq)
    elif method == "black_litterman":
        return black_litterman_weights(prices=prices, rfr=rfr, freq=freq, **kwargs)
    elif method == "equal":
        return equal_weighted(prices=prices, returns=returns)
    else:
        raise ValueError(f"Unknown method: {method}")

nlp

SentimentAnalyzer

SentimentAnalyzer()

Bases: object

A financial sentiment analysis tool that processes and analyzes sentiment from news articles, social media posts, and financial reports.

This class utilizes NLP techniques to preprocess text and apply sentiment analysis using VADER (SentimentIntensityAnalyzer) and optional TextBlob for enhanced polarity scoring.

Initializes the SentimentAnalyzer class by downloading necessary NLTK resources and loading the SpaCy NLP model.

  • Downloads NLTK tokenization (punkt) and stopwords.
  • Loads the en_core_web_sm SpaCy model with Named Entity Recognition (NER) disabled.
  • Initializes VADER's SentimentIntensityAnalyzer for sentiment scoring.
Source code in src/bbstrader/models/nlp.py
def __init__(self):
    """
    Initializes the SentimentAnalyzer class by downloading necessary
    NLTK resources and loading the SpaCy NLP model.

    - Downloads NLTK tokenization (`punkt`) and stopwords.
    - Loads the `en_core_web_sm` SpaCy model with Named Entity Recognition (NER) disabled.
    - Initializes VADER's SentimentIntensityAnalyzer for sentiment scoring.

    """
    _require_nlp()
    nltk.download("punkt", quiet=True)
    nltk.download("punkt_tab", quiet=True)
    nltk.download("stopwords", quiet=True)

    self.analyzer = SentimentIntensityAnalyzer()
    self._stopwords = set(stopwords.words("english"))

    try:
        self.nlp = spacy.load("en_core_web_sm", disable=["ner"])
    except OSError:
        print("Downloading 'en_core_web_sm' model for spaCy...")
        download("en_core_web_sm")
        self.nlp = spacy.load("en_core_web_sm", disable=["ner"])

    self.news = FinancialNews()
preprocess_text
preprocess_text(text: str)

Preprocesses the input text by performing the following steps: 1. Converts text to lowercase. 2. Removes URLs. 3. Removes all non-alphabetic characters (punctuation, numbers, special symbols). 4. Tokenizes the text into words. 5. Removes stop words. 6. Lemmatizes the words using SpaCy, excluding pronouns.

Parameters:

Name Type Description Default
text str

The input text to preprocess.

required

Returns:

Name Type Description
str

The cleaned and lemmatized text.

Source code in src/bbstrader/models/nlp.py
def preprocess_text(self, text: str):
    """
    Preprocesses the input text by performing the following steps:
    1. Converts text to lowercase.
    2. Removes URLs.
    3. Removes all non-alphabetic characters (punctuation, numbers, special symbols).
    4. Tokenizes the text into words.
    5. Removes stop words.
    6. Lemmatizes the words using SpaCy, excluding pronouns.

    Args:
        text (str): The input text to preprocess.

    Returns:
        str: The cleaned and lemmatized text.
    """
    if not isinstance(text, str):
        raise ValueError(
            f"{self.__class__.__name__}: preprocess_text expects a string, got {type(text)}"
        )
    text = text.lower()
    text = re.sub(r"http\S+", "", text)
    text = re.sub(r"[^a-zA-Z\s]", "", text)

    words = word_tokenize(text)
    words = [word for word in words if word not in self._stopwords]

    doc = self.nlp(" ".join(words))
    words = [t.lemma_ for t in doc if t.lemma_ != "-PRON-"]

    return " ".join(words)
analyze_sentiment
analyze_sentiment(texts, lexicon=None, textblob=False) -> float

Analyzes the sentiment of a list of texts using VADER or TextBlob.

Steps: 1. If a custom lexicon is provided, updates the VADER lexicon. 2. If textblob is set to True, computes sentiment using TextBlob. 3. Otherwise, preprocesses the text and computes sentiment using VADER. 4. Returns the average sentiment score of all input texts.

Parameters:

Name Type Description Default
texts list of str

A list of text inputs to analyze.

required
lexicon dict

A custom sentiment lexicon to update VADER's default lexicon.

None
textblob bool

If True, uses TextBlob for sentiment analysis instead of VADER.

False

Returns:

Name Type Description
float float

The average sentiment score across all input texts. - Positive values indicate positive sentiment. - Negative values indicate negative sentiment. - Zero indicates neutral sentiment.

Source code in src/bbstrader/models/nlp.py
def analyze_sentiment(self, texts, lexicon=None, textblob=False) -> float:
    """
    Analyzes the sentiment of a list of texts using VADER or TextBlob.

    Steps:
    1. If a custom lexicon is provided, updates the VADER lexicon.
    2. If `textblob` is set to True, computes sentiment using TextBlob.
    3. Otherwise, preprocesses the text and computes sentiment using VADER.
    4. Returns the average sentiment score of all input texts.

    Args:
        texts (list of str): A list of text inputs to analyze.
        lexicon (dict, optional): A custom sentiment lexicon to update VADER's default lexicon.
        textblob (bool, optional): If True, uses TextBlob for sentiment analysis instead of VADER.

    Returns:
        float: The average sentiment score across all input texts.
               - Positive values indicate positive sentiment.
               - Negative values indicate negative sentiment.
               - Zero indicates neutral sentiment.
    """
    if lexicon is not None:
        self.analyzer.lexicon.update(lexicon)
    if textblob:
        blob = TextBlob(" ".join(texts))
        return blob.sentiment.polarity
    sentiment_scores = [
        self.analyzer.polarity_scores(self.preprocess_text(text))["compound"]
        for text in texts
    ]
    avg_sentiment = (
        sum(sentiment_scores) / len(sentiment_scores) if sentiment_scores else 0.0
    )
    return avg_sentiment
get_sentiment_for_tickers
get_sentiment_for_tickers(tickers: List[str] | List[Tuple[str, str]], lexicon=None, asset_type='stock', top_news=10, **kwargs) -> Dict[str, float]

Compute sentiment scores for a list of financial tickers based on news and social media data.

Process
  1. Collect news articles and posts related to each ticker from various sources:
  2. Yahoo Finance News
  3. Google Finance News
  4. Reddit posts
  5. Financial Modeling Prep (FMP) news
  6. Analyze sentiment from each source:
  7. Uses VADER for Yahoo and Google Finance news.
  8. Uses TextBlob for Reddit and FMP news.
  9. Compute an overall sentiment score using a weighted average approach.
Parameters

tickers : list of str or list of tuple A list of asset tickers to analyze. * If using tuples, the first element is the ticker and the second is the asset type. * If using a single string, the asset type must be specified or defaults to "stock". lexicon : dict, optional A custom sentiment lexicon to update VADER's default lexicon. Default is None. asset_type : str, optional The type of asset. Default is "stock". Supported types include: * "stock": Stock symbols (e.g., AAPL, MSFT) * "etf": Exchange-traded funds (e.g., SPY, QQQ) * "future": Futures contracts (e.g., CL=F for crude oil) * "forex": Forex pairs (e.g., EURUSD=X, USDJPY=X) * "crypto": Cryptocurrency pairs (e.g., BTC-USD, ETH-USD) * "index": Stock market indices (e.g., ^GSPC for S&P 500) top_news : int, optional Number of news articles/posts to fetch per source. Default is 10. **kwargs : dict Additional parameters for API authentication and data retrieval. Must include: * fmp_api (str): API key for Financial Modeling Prep. * client_id, client_secret, user_agent (str): Credentials for Reddit API.

Returns

dict of str to float A dictionary mapping each ticker to its overall sentiment score. * Positive values indicate positive sentiment. * Negative values indicate negative sentiment. * Zero indicates neutral sentiment.

Notes

Ticker names must follow Yahoo Finance conventions.

Source code in src/bbstrader/models/nlp.py
def get_sentiment_for_tickers(
    self,
    tickers: List[str] | List[Tuple[str, str]],
    lexicon=None,
    asset_type="stock",
    top_news=10,
    **kwargs,
) -> Dict[str, float]:
    """
    Compute sentiment scores for a list of financial tickers based on news and social media data.

    Process
    -------
    1. Collect news articles and posts related to each ticker from various sources:
    * Yahoo Finance News
    * Google Finance News
    * Reddit posts
    * Financial Modeling Prep (FMP) news
    2. Analyze sentiment from each source:
    * Uses VADER for Yahoo and Google Finance news.
    * Uses TextBlob for Reddit and FMP news.
    3. Compute an overall sentiment score using a weighted average approach.

    Parameters
    ----------
    tickers : list of str or list of tuple
        A list of asset tickers to analyze.
        * If using tuples, the first element is the ticker and the second is the asset type.
        * If using a single string, the asset type must be specified or defaults to "stock".
    lexicon : dict, optional
        A custom sentiment lexicon to update VADER's default lexicon. Default is None.
    asset_type : str, optional
        The type of asset. Default is "stock".
        Supported types include:
        * "stock": Stock symbols (e.g., AAPL, MSFT)
        * "etf": Exchange-traded funds (e.g., SPY, QQQ)
        * "future": Futures contracts (e.g., CL=F for crude oil)
        * "forex": Forex pairs (e.g., EURUSD=X, USDJPY=X)
        * "crypto": Cryptocurrency pairs (e.g., BTC-USD, ETH-USD)
        * "index": Stock market indices (e.g., ^GSPC for S&P 500)
    top_news : int, optional
        Number of news articles/posts to fetch per source. Default is 10.
    **kwargs : dict
        Additional parameters for API authentication and data retrieval. Must include:
        * fmp_api (str): API key for Financial Modeling Prep.
        * client_id, client_secret, user_agent (str): Credentials for Reddit API.

    Returns
    -------
    dict of str to float
        A dictionary mapping each ticker to its overall sentiment score.
        * Positive values indicate positive sentiment.
        * Negative values indicate negative sentiment.
        * Zero indicates neutral sentiment.

    Notes
    -----
    Ticker names must follow Yahoo Finance conventions.
    """

    sentiment_results = {}

    # Suppress stdout/stderr from underlying  libraries during execution
    with open(os.devnull, "w") as devnull:
        with (
            contextlib.redirect_stdout(devnull),
            contextlib.redirect_stderr(devnull),
        ):
            with ThreadPoolExecutor() as executor:
                # Map each future to its ticker for easy result lookup
                future_to_ticker = {}
                for ticker_info in tickers:
                    # Normalize input to (ticker, asset_type)
                    if isinstance(ticker_info, tuple):
                        ticker_symbol, ticker_asset_type = ticker_info
                    else:
                        ticker_symbol, ticker_asset_type = ticker_info, asset_type

                    if ticker_asset_type not in [
                        "stock",
                        "etf",
                        "future",
                        "forex",
                        "crypto",
                        "index",
                    ]:
                        raise ValueError(
                            f"Unsupported asset type '{ticker_asset_type}' for {ticker_symbol}."
                        )

                    # Submit the job to the thread pool
                    future = executor.submit(
                        self._get_sentiment_for_one_ticker,
                        ticker=ticker_symbol,
                        asset_type=ticker_asset_type,
                        lexicon=lexicon,
                        top_news=top_news,
                        **kwargs,
                    )
                    future_to_ticker[future] = ticker_symbol

                # Collect results as they are completed
                for future in as_completed(future_to_ticker):
                    ticker_symbol = future_to_ticker[future]
                    try:
                        sentiment_score = future.result()
                        sentiment_results[ticker_symbol] = sentiment_score
                    except Exception:
                        sentiment_results[ticker_symbol] = (
                            0.0  # Assign a neutral score on error
                        )

    return sentiment_results
get_topn_sentiments
get_topn_sentiments(sentiments, topn=10)

Retrieves the top and bottom N assets based on sentiment scores.

Parameters:

Name Type Description Default
sentiments dict

A dictionary mapping asset tickers to their sentiment scores.

required
topn int

The number of top and bottom assets to return. Defaults to 10.

10

Returns:

Name Type Description
tuple

A tuple containing two lists: - bottom (list of tuples): The topn assets with the lowest sentiment scores, sorted in ascending order. - top (list of tuples): The topn assets with the highest sentiment scores, sorted in descending order.

Source code in src/bbstrader/models/nlp.py
def get_topn_sentiments(self, sentiments, topn=10):
    """
    Retrieves the top and bottom N assets based on sentiment scores.

    Args:
        sentiments (dict): A dictionary mapping asset tickers to their sentiment scores.
        topn (int, optional): The number of top and bottom assets to return. Defaults to 10.

    Returns:
        tuple: A tuple containing two lists:
            - bottom (list of tuples): The `topn` assets with the lowest sentiment scores, sorted in ascending order.
            - top (list of tuples): The `topn` assets with the highest sentiment scores, sorted in descending order.
    """
    sorted_sentiments = sorted(sentiments.items(), key=lambda x: x[1])
    bottom = sorted_sentiments[:topn]
    top = sorted_sentiments[-topn:]
    return bottom, top
visualize_sentiments
visualize_sentiments(sentiment_dict, mode='bar', top_n=10)

Visualizes sentiment scores for financial assets using different chart types.

Visualization Modes: - "bar": Displays a bar chart of the top N assets by sentiment score. - "scatter": Displays a scatter plot of sentiment scores.

Parameters:

Name Type Description Default
sentiment_dict dict

A dictionary mapping asset tickers to their sentiment scores.

required
mode str

The type of visualization to generate. Options: "bar" (default), "scatter".

'bar'
top_n int

The number of top tickers to display in the bar chart. Only applicable when mode is "bar".

10

Returns:

Name Type Description
None

Displays the sentiment visualization.

Source code in src/bbstrader/models/nlp.py
def visualize_sentiments(self, sentiment_dict, mode="bar", top_n=10):
    """
    Visualizes sentiment scores for financial assets using different chart types.

    Visualization Modes:
    - "bar": Displays a bar chart of the top N assets by sentiment score.
    - "scatter": Displays a scatter plot of sentiment scores.

    Args:
        sentiment_dict (dict): A dictionary mapping asset tickers to their sentiment scores.
        mode (str, optional): The type of visualization to generate.
                              Options: "bar" (default), "scatter".
        top_n (int, optional): The number of top tickers to display in the bar chart.
                               Only applicable when mode is "bar".

    Returns:
        None: Displays the sentiment visualization.
    """
    if mode == "bar":
        self._sentiment_bar(sentiment_dict, top_n=top_n)
    elif mode == "scatter":
        self._sentiment_scatter(sentiment_dict)

optimization

markowitz_weights

markowitz_weights(prices=None, rfr=0.0, freq=252, min_vol=False)

Calculates optimal portfolio weights using Markowitz's mean-variance optimization (Max Sharpe Ratio or Min Volatility) with multiple solvers.

Parameters

prices : pd.DataFrame, optional Price data for assets, where rows represent time periods and columns represent assets. rfr : float, optional Risk-free rate (default is 0.0). freq : int, optional Frequency of the data, such as 252 for daily returns in a year (default is 252). min_vol : bool, optional If True, optimizes for minimum volatility instead of maximum Sharpe ratio (default is False).

Returns

dict Dictionary containing the optimal asset weights for maximizing the Sharpe ratio or minimizing volatility, normalized to sum to 1.

Notes

This function attempts to maximize the Sharpe ratio by iterating through various solvers ('SCS', 'ECOS', 'OSQP') from the PyPortfolioOpt library. If a solver fails, it proceeds to the next one. If none succeed, an error message is printed for each solver that fails.

This function is useful for portfolio with a small number of assets, as it may not scale well for large portfolios.

Raises

Exception If all solvers fail, each will print an exception error message during runtime.

Source code in src/bbstrader/models/optimization.py
def markowitz_weights(prices=None, rfr=0.0, freq=252, min_vol=False):
    """
    Calculates optimal portfolio weights using Markowitz's mean-variance optimization (Max Sharpe Ratio or Min Volatility) with multiple solvers.

    Parameters
    ----------
    prices : pd.DataFrame, optional
        Price data for assets, where rows represent time periods and columns represent assets.
    rfr : float, optional
        Risk-free rate (default is 0.0).
    freq : int, optional
        Frequency of the data, such as 252 for daily returns in a year (default is 252).
    min_vol : bool, optional
        If True, optimizes for minimum volatility instead of maximum Sharpe ratio (default is False).

    Returns
    -------
    dict
        Dictionary containing the optimal asset weights for maximizing the Sharpe ratio or minimizing volatility, normalized to sum to 1.

    Notes
    -----
    This function attempts to maximize the Sharpe ratio by iterating through various solvers ('SCS', 'ECOS', 'OSQP')
    from the PyPortfolioOpt library. If a solver fails, it proceeds to the next one. If none succeed, an error message
    is printed for each solver that fails.

    This function is useful for portfolio with a small number of assets, as it may not scale well for large portfolios.

    Raises
    ------
    Exception
        If all solvers fail, each will print an exception error message during runtime.
    """
    returns = expected_returns.mean_historical_return(prices, frequency=freq)
    cov = risk_models.sample_cov(prices, frequency=freq)

    # Try different solvers to maximize Sharpe ratio
    for solver in ["SCS", "ECOS", "OSQP"]:
        ef = EfficientFrontier(
            expected_returns=returns,
            cov_matrix=cov,
            weight_bounds=(0, 1),
            solver=solver,
        )
        try:
            if min_vol:
                ef.min_volatility()
            else:
                ef.max_sharpe(risk_free_rate=rfr)
            return _normalize_weights(ef.clean_weights())
        except Exception as e:
            print(f"Solver {solver} failed with error: {e}")
    # Default to equal weighted if all solvers fail
    return _normalize_weights(equal_weighted(prices=prices))

hierarchical_risk_parity

hierarchical_risk_parity(prices=None, returns=None, freq=252)

Computes asset weights using Hierarchical Risk Parity (HRP) for risk-averse portfolio allocation.

Parameters

prices : pd.DataFrame, optional Price data for assets; if provided, daily returns will be calculated. returns : pd.DataFrame, optional Daily returns for assets. One of prices or returns must be provided. freq : int, optional Number of days to consider in calculating portfolio weights (default is 252).

Returns

dict Optimized asset weights using the HRP method, with asset weights summing to 1.

Raises

ValueError If neither prices nor returns are provided.

Notes

Hierarchical Risk Parity is particularly useful for portfolios with a large number of assets, as it mitigates issues of multicollinearity and estimation errors in covariance matrices by using hierarchical clustering.

Source code in src/bbstrader/models/optimization.py
def hierarchical_risk_parity(prices=None, returns=None, freq=252):
    """
    Computes asset weights using Hierarchical Risk Parity (HRP) for risk-averse portfolio allocation.

    Parameters
    ----------
    prices : pd.DataFrame, optional
        Price data for assets; if provided, daily returns will be calculated.
    returns : pd.DataFrame, optional
        Daily returns for assets. One of `prices` or `returns` must be provided.
    freq : int, optional
        Number of days to consider in calculating portfolio weights (default is 252).

    Returns
    -------
    dict
        Optimized asset weights using the HRP method, with asset weights summing to 1.

    Raises
    ------
    ValueError
        If neither `prices` nor `returns` are provided.

    Notes
    -----
    Hierarchical Risk Parity is particularly useful for portfolios with a large number of assets,
    as it mitigates issues of multicollinearity and estimation errors in covariance matrices by
    using hierarchical clustering.
    """
    warnings.filterwarnings("ignore")
    if returns is None and prices is None:
        raise ValueError("Either prices or returns must be provided")
    if returns is None:
        returns = prices.pct_change().dropna(how="all")
    # Remove duplicate columns and index
    returns = returns.loc[:, ~returns.columns.duplicated()]
    returns = returns.loc[~returns.index.duplicated(keep="first")]
    hrp = HRPOpt(returns=returns.iloc[-freq:])
    return _normalize_weights(hrp.optimize())

equal_weighted

equal_weighted(prices=None, returns=None, round_digits=5)

Generates an equal-weighted portfolio by assigning an equal proportion to each asset.

Parameters

prices : pd.DataFrame, optional Price data for assets, where each column represents an asset. returns : pd.DataFrame, optional Return data for assets. One of prices or returns must be provided. round_digits : int, optional Number of decimal places to round each weight to (default is 5).

Returns

dict Dictionary with equal weights assigned to each asset, summing to 1.

Raises

ValueError If neither prices nor returns are provided.

Notes

Equal weighting is a simple allocation method that assumes equal importance across all assets, useful as a baseline model and when no strong views exist on asset return expectations or risk.

Source code in src/bbstrader/models/optimization.py
def equal_weighted(prices=None, returns=None, round_digits=5):
    """
    Generates an equal-weighted portfolio by assigning an equal proportion to each asset.

    Parameters
    ----------
    prices : pd.DataFrame, optional
        Price data for assets, where each column represents an asset.
    returns : pd.DataFrame, optional
        Return data for assets. One of `prices` or `returns` must be provided.
    round_digits : int, optional
        Number of decimal places to round each weight to (default is 5).

    Returns
    -------
    dict
        Dictionary with equal weights assigned to each asset, summing to 1.

    Raises
    ------
    ValueError
        If neither `prices` nor `returns` are provided.

    Notes
    -----
    Equal weighting is a simple allocation method that assumes equal importance across all assets,
    useful as a baseline model and when no strong views exist on asset return expectations or risk.
    """

    if returns is None and prices is None:
        raise ValueError("Either prices or returns must be provided")
    if returns is None:
        n = len(prices.columns)
        columns = prices.columns
    else:
        n = len(returns.columns)
        columns = returns.columns
    return {col: round(1 / n, round_digits) for col in columns}

black_litterman_weights

black_litterman_weights(prices=None, rfr=0.0, freq=252, views=None, view_confidences=None, pi=None, market_caps=None)

Computes portfolio weights using the Black-Litterman model.

Parameters

prices : pd.DataFrame Price data for assets. rfr : float, optional Risk-free rate (default is 0.0). freq : int, optional Frequency of the data (default is 252). views : dict, optional Investor's views on asset returns. view_confidences : list or np.array, optional Confidence levels for each view. pi : pd.Series, optional Market-implied prior returns. market_caps : pd.Series, optional Market capitalization of assets.

Returns

dict Optimal asset weights based on the Black-Litterman model.

Source code in src/bbstrader/models/optimization.py
def black_litterman_weights(
    prices=None,
    rfr=0.0,
    freq=252,
    views=None,
    view_confidences=None,
    pi=None,
    market_caps=None,
):
    """
    Computes portfolio weights using the Black-Litterman model.

    Parameters
    ----------
    prices : pd.DataFrame
        Price data for assets.
    rfr : float, optional
        Risk-free rate (default is 0.0).
    freq : int, optional
        Frequency of the data (default is 252).
    views : dict, optional
        Investor's views on asset returns.
    view_confidences : list or np.array, optional
        Confidence levels for each view.
    pi : pd.Series, optional
        Market-implied prior returns.
    market_caps : pd.Series, optional
        Market capitalization of assets.

    Returns
    -------
    dict
        Optimal asset weights based on the Black-Litterman model.
    """
    cov_matrix = risk_models.sample_cov(prices, frequency=freq)
    if pi is None:
        if market_caps is not None:
            # If market caps are provided, we can use them to compute the prior
            # This requires a benchmark, which we don't have here easily.
            # For simplicity, we use the mean historical return if pi is not provided.
            pi = expected_returns.mean_historical_return(prices, frequency=freq)
        else:
            pi = expected_returns.mean_historical_return(prices, frequency=freq)

    bl = BlackLittermanModel(
        cov_matrix,
        pi=pi,
        absolute_views=views,
        omega=None,
        view_confidences=view_confidences,
    )
    ret_bl = bl.bl_returns()
    ef = EfficientFrontier(ret_bl, cov_matrix)
    ef.max_sharpe(risk_free_rate=rfr)
    return _normalize_weights(ef.clean_weights())

optimized_weights

optimized_weights(prices=None, returns=None, rfr=0.0, freq=252, method='equal', **kwargs)

Selects an optimization method to calculate portfolio weights based on user preference.

Parameters

prices : pd.DataFrame, optional Price data for assets, required for certain methods. returns : pd.DataFrame, optional Returns data for assets, an alternative input for certain methods. freq : int, optional Number of days for calculating portfolio weights, such as 252 for a year's worth of daily returns (default is 252). method : str, optional Optimization method to use ('markowitz', 'hrp', or 'equal') (default is 'equal').

Returns

dict Dictionary containing optimized asset weights based on the chosen method.

Raises

ValueError If an unknown optimization method is specified.

Notes

This function integrates different optimization methods: - 'markowitz': mean-variance optimization with max Sharpe ratio - 'hrp': Hierarchical Risk Parity, for risk-based clustering of assets - 'equal': Equal weighting across all assets

Source code in src/bbstrader/models/optimization.py
def optimized_weights(
    prices=None, returns=None, rfr=0.0, freq=252, method="equal", **kwargs
):
    """
    Selects an optimization method to calculate portfolio weights based on user preference.

    Parameters
    ----------
    prices : pd.DataFrame, optional
        Price data for assets, required for certain methods.
    returns : pd.DataFrame, optional
        Returns data for assets, an alternative input for certain methods.
    freq : int, optional
        Number of days for calculating portfolio weights, such as 252 for a year's worth of daily returns (default is 252).
    method : str, optional
        Optimization method to use ('markowitz', 'hrp', or 'equal') (default is 'equal').

    Returns
    -------
    dict
        Dictionary containing optimized asset weights based on the chosen method.

    Raises
    ------
    ValueError
        If an unknown optimization method is specified.

    Notes
    -----
    This function integrates different optimization methods:
    - 'markowitz': mean-variance optimization with max Sharpe ratio
    - 'hrp': Hierarchical Risk Parity, for risk-based clustering of assets
    - 'equal': Equal weighting across all assets
    """
    if method == "markowitz":
        return markowitz_weights(prices=prices, rfr=rfr, freq=freq)
    elif method == "min_vol":
        return markowitz_weights(prices=prices, rfr=rfr, freq=freq, min_vol=True)
    elif method == "hrp":
        return hierarchical_risk_parity(prices=prices, returns=returns, freq=freq)
    elif method == "black_litterman":
        return black_litterman_weights(prices=prices, rfr=rfr, freq=freq, **kwargs)
    elif method == "equal":
        return equal_weighted(prices=prices, returns=returns)
    else:
        raise ValueError(f"Unknown method: {method}")